spot_img
HomeResearch & DevelopmentChart2Code: A New Benchmark Reveals Gaps in AI's Chart...

Chart2Code: A New Benchmark Reveals Gaps in AI’s Chart Generation Abilities

TLDR: Chart2Code is a new hierarchical benchmark evaluating large multimodal models (LMMs) on chart understanding and code generation, designed to reflect real-world usage. It features three levels of increasing difficulty: chart reproduction, editing, and long-table to chart generation. Benchmarking 25 state-of-the-art LMMs, the study reveals that even top models struggle significantly with complex editing and long-context data-to-chart tasks, particularly in achieving visual fidelity despite generating executable code. The benchmark aims to drive advancements in robust LMMs for visualization.

A new research paper introduces Chart2Code, a groundbreaking benchmark designed to thoroughly evaluate how well large multimodal models (LMMs) can understand charts and generate the code needed to create them. This benchmark is unique because it’s built from a user’s perspective, reflecting real-world challenges and gradually increasing in difficulty.

What is Chart2Code?

Chart2Code is structured into three distinct levels, each presenting a different degree of complexity for LMMs:

Level 1: Chart Reproduction involves tasks where models must either perfectly recreate a chart from a reference image and a user’s request, or extract data from a file and then generate a chart that matches the style of a given reference. This level tests a model’s basic visual understanding and its ability to translate visual information into executable code.

Level 2: Chart Editing steps up the challenge by requiring models to make complex modifications to existing charts. This could mean changing the chart type, merging different datasets, or adding new visual elements, all while following specific user instructions. It assesses a model’s capacity for precise and nuanced chart manipulation.

Level 3: Long-Table to Chart Generation is the most demanding level. Here, models are given long, information-dense tables and must transform this raw data into accurate and faithful charts based on user instructions and a reference style. This level simulates real-world scenarios where users are not data visualization experts and need AI to distill complex data into clear visuals.

Why is Chart2Code Important?

Existing benchmarks often suggest that LMMs are already highly capable at chart-to-code tasks. However, in practical, everyday scenarios, these same models often fall short. Chart2Code addresses this gap by providing a benchmark that truly reflects how people use charts – not just for simple reproduction, but for editing, customizing, and generating from complex raw data. It includes 2,023 tasks across 22 chart types, with multi-level evaluation metrics that check both the correctness of the generated code and the visual accuracy of the resulting charts.

Key Findings

The researchers benchmarked 25 state-of-the-art LMMs, including proprietary models like GPT-5 and open-source models such as Qwen2.5-VL and InternVL3/3.5. The results highlight significant challenges: even the most advanced models, like GPT-5, achieved an average of only 0.57 on code-based evaluation and 0.22 on chart-quality assessment for editing tasks. This demonstrates that while models might generate code that runs, achieving pixel-perfect visual fidelity remains a major hurdle. Performance drops sharply as task complexity increases, especially with long-context data-to-chart generation. The study also found that while “thinking” or chain-of-thought approaches can improve a model’s ability to follow instructions and execute code, these improvements don’t consistently translate to better visual fidelity.

Also Read:

Looking Ahead

Chart2Code is expected to be a crucial tool for driving advancements in multimodal reasoning and fostering the development of more robust and versatile LMMs. The benchmark’s code and data are publicly available, encouraging further research in this area. For more details, you can read the full research paper: FROM CHARTS TO CODE: A HIERARCHICAL BENCHMARK FOR MULTIMODAL MODELS.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -