TLDR: A new benchmark, GSM8K-V, converts text-based grade school math word problems into purely visual, multi-image formats to evaluate Vision Language Models (VLMs). It reveals a significant performance gap, with VLMs struggling much more on visual math (e.g., Gemini-2.5-Pro at 46.93%) compared to text (95.22%), highlighting critical limitations in visual perception, instrument reading, and multi-image reasoning.
Vision Language Models (VLMs) have made incredible strides in understanding both images and text, allowing them to tackle complex tasks that involve perception, planning, and reasoning. Among these, mathematical reasoning stands out as a particularly challenging area, especially when the math problems are presented in a visual format.
While many benchmarks exist to test VLMs’ mathematical abilities, they often fall short. They tend to focus on geometry problems, don’t fully evaluate math word problems in visual contexts, and rarely assess reasoning across multiple images. This leaves a significant gap in understanding how well VLMs can truly solve real-world math problems when information is presented visually.
Introducing GSM8K-V: A New Visual Math Benchmark
To address these limitations, researchers have introduced GSM8K-V, a novel benchmark designed specifically for purely visual, multi-image mathematical reasoning. This benchmark takes the widely used text-based math reasoning problems from GSM8K and transforms them into a comic-style visual format.
The creation of GSM8K-V involved a sophisticated automated pipeline combined with careful human annotation. This process systematically converted each text problem into a series of visual scenes, resulting in a high-quality dataset of 1,319 samples. The pipeline ensures that mathematical information is accurately extracted, classified, and then allocated across different scenes, sometimes even introducing controlled visual or semantic distractions to increase difficulty.
Key Findings: A Significant Performance Gap
The evaluation of various open-source and closed-source VLMs on GSM8K-V revealed a striking disparity in performance. While existing VLMs achieve nearly perfect scores (often 80-90% accuracy) on the original text-based GSM8K, their performance drops dramatically on the purely visual GSM8K-V.
For instance, the top-performing model, Gemini-2.5-Pro, achieved an impressive 95.22% accuracy on text-based GSM8K but only managed 46.93% on GSM8K-V. Other leading models also showed similar struggles, with many scoring around 30%. This stark contrast highlights that current VLMs face substantial challenges in interpreting and reasoning over mathematical information embedded in images, especially when multiple images are involved.
Interestingly, human performance on GSM8K-V is much higher, averaging 91.15% accuracy, and humans maintain consistent performance across different problem categories. VLMs, however, show significant variability, performing better on some categories like ‘Signboard & Icon’ but struggling greatly with others, suggesting a lack of generalized visual reasoning.
Understanding the Challenges
The research identified two primary types of errors contributing to VLMs’ underperformance:
- Perception-Calculation Errors: Models often misinterpret visual details, such as miscounting objects or confusing similar items, leading to incorrect calculations. Examples include misinterpreting flower package sizes or lumber board counts.
- Instrument-Reading Errors: VLMs struggle to accurately read and interpret common instruments like clocks, gauges, or pie charts that encode numerical values. This can lead to fundamental errors in reasoning, even for straightforward arithmetic problems.
Further analysis showed that providing explicit textual questions alongside images only slightly improved performance, and concatenating multiple scenes into a single image often obscured critical temporal and logical relationships, making problems harder. This confirms that GSM8K-V genuinely requires visual reasoning, as models cannot simply rely on text transcriptions (OCR) of the images.
Also Read:
- Unveiling VLM Limitations in Visually Complex Environments
- Why Advanced AI Models Struggle with Simple Visual Tasks: The Serial Processing Gap
Future Directions for VLM Development
GSM8K-V offers a crucial new perspective on visual mathematical reasoning. By providing a reliable and challenging benchmark, it guides the research community toward developing more robust and generalizable VLMs that can truly understand and solve math word problems in diverse visual contexts. The significant performance gap revealed by this benchmark underscores the need for continued innovation in multimodal AI to bridge the divide between textual and visual reasoning capabilities.


