spot_img
HomeResearch & DevelopmentAdvancing AI's Geometric Reasoning with a New Data Synthesis...

Advancing AI’s Geometric Reasoning with a New Data Synthesis Framework

TLDR: A new research paper introduces Geo-Image-Textualization, a reinforcement learning-based framework that creates high-quality, fully aligned geometric image-caption pairs. This framework generates GeoReasoning-10K, a dataset that significantly improves Multimodal Large Language Models’ (MLLMs) ability to solve geometric problems and enhances their general reasoning capabilities across various mathematical and non-mathematical domains, even with non-geometric inputs.

Multimodal Large Language Models, or MLLMs, have shown impressive abilities in various tasks that involve both images and text, such as answering questions about visuals and describing images. However, these advanced AI models still face significant challenges when it comes to solving complex geometric problems. A major reason for this difficulty is the scarcity of high-quality datasets that pair geometric images with accurate textual descriptions, which are crucial for training MLLMs to understand geometry.

Existing methods for creating such datasets often fall short because their templates don’t generalize well to new or different types of geometric questions. This means the models trained on these datasets struggle to apply their knowledge beyond the specific examples they were shown.

Researchers have introduced a new approach to tackle this problem, integrating a process called Reinforcement Learning with Verifiable Rewards (RLVR) into the data generation pipeline. This method refines captions for geometric images that are synthesized from 50 basic geometric relationships. By using reward signals derived from mathematical problem-solving tasks, the pipeline effectively captures the essential features of geometry problem-solving. This leads to better generalization across tasks and noticeable improvements in performance.

The new pipeline, named Geo-Image-Textualization, is an RL-based data refinement engine that continuously improves data quality. Using this framework, the team developed a novel geometry dataset called GeoReasoning-10K, which contains 10,000 image-caption pairs. What makes this dataset unique is that the visual and textual information are fully aligned, and it generalizes well to tasks outside its original distribution. This makes it a valuable resource for enhancing cross-modal reasoning in MLLMs.

The Geo-Image-Textualization pipeline involves three main parts: relation sampling, image-caption pair generation, and question-answer generation. The relation sampling uses fundamental geometric operations to create diverse and coherent geometric premises. The image-caption pair generation addresses a key limitation of previous methods by explicitly encoding semantic relationships within the images using visual augmentation strategies. These strategies include using ticks for equal-length segments, arcs for equal angles, and symbols for parallel or perpendicular lines, ensuring that all visual information is accurately reflected in the captions.

The question-answer pair generation uses a rule-and-LLM-based pipeline, prompting a large language model (Gemini 2.5 Flash) to create questions based on the captions, while ensuring consistency and avoiding repetition of information already in the caption.

The RLVR framework then iteratively optimizes both the model and the dataset. It starts with a “cold-start” supervised fine-tuning phase to give the model initial captioning abilities using the GeoReasoning-10K dataset. Following this, the RLVR phase uses a technique called RAFT (Reward Ranked Finetuning) to cyclically refine the dataset and model through reinforcement learning. A composite reward function balances task correctness and caption-image alignment, using both reasoning rewards (evaluating utility for solving tasks) and caption rewards (measuring semantic relevance to ground-truth captions).

Experiments using Gemma3-4B as the base model showed that models trained on GeoReasoning-10K achieved better mathematical reasoning performance on benchmarks like MathVista and MathVerse compared to models trained on other datasets. This improvement was seen across various mathematical subtasks, including geometry, algebra, science, and statistics. The dataset also demonstrated better scalability, with performance improving as dataset size increased.

Remarkably, GeoReasoning-10K also showed improved generalization abilities for non-geometric input images. Models trained on this dataset outperformed baselines on the MMMU benchmark, particularly in domains like Art & Design and Tech & Engineering. This suggests that the RLVR training process, by forcing the model to focus on key elements for problem-solving, helps it generalize to scenarios beyond geometric problems.

Also Read:

Ablation studies confirmed that both the cold-start and RLVR phases are effective, with RLVR being particularly helpful. The research highlights the potential of rule-based symbolic synthesis for cross-modal and cross-domain learning, offering a promising direction for future research in multimodal reasoning. For more details, you can read the full research paper here.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -