spot_img
HomeResearch & DevelopmentNew Benchmark Reveals AI Strengths and Weaknesses in Scatterplot...

New Benchmark Reveals AI Strengths and Weaknesses in Scatterplot Analysis

TLDR: Researchers have introduced the “Benchmark It Yourself” (BIY) dataset and benchmark, featuring over 18,000 synthetic scatterplots, to evaluate AI models on scatterplot-specific tasks. The study tested OpenAI and Google’s Gemini models on tasks like counting clusters and identifying outliers. While models showed high accuracy (90%+) in counting tasks with few-shot prompting, their performance in precise localization (detection and identification) was generally unsatisfactory. The research also highlighted that wide aspect ratios and randomly colored charts can negatively impact AI performance, recommending few-shot prompting and caution for localization tasks.

Artificial intelligence models are increasingly being used for data analysis and visualization, but a significant gap exists in how well these models perform on tasks specifically related to scatterplots. To address this, a new research paper introduces a comprehensive benchmark and dataset called “Benchmark It Yourself” (BIY).

The BIY initiative provides a synthetic, annotated dataset comprising over 18,000 scatterplots. These plots are generated from six different data generators and feature 17 distinct chart designs, making the dataset robust and varied. The goal is to offer a clearer understanding of the actual capabilities of current AI models when interpreting visual scatterplot data.

The researchers evaluated ten proprietary AI models from OpenAI (including GPT-4.1, GPT-4o, and their mini/nano variants) and Google (Gemini 2.5 Flash and Flash-Lite). These models were tested on five distinct scatterplot-related tasks:

The Five Key Tasks

  • Cluster counting: Determining the number of clusters present.
  • Cluster detection: Identifying the bounding boxes around clusters.
  • Cluster identification: Pinpointing the center coordinates of clusters.
  • Outlier counting: Counting the number of anomalous data points.
  • Outlier identification: Locating the precise coordinates of outliers.

Three prompting strategies were employed: zero-shot (no examples), one-shot (one example), and few-shot (multiple examples). The findings highlight some interesting performance trends. For counting tasks, such as counting clusters and outliers, OpenAI models and Gemini 2.5 Flash demonstrated strong performance, achieving over 90% accuracy, especially when provided with a few examples (few-shot prompting).

However, the results for localization-related tasks—like precisely detecting cluster bounding boxes or identifying exact outlier coordinates—were less satisfactory. Precision and Recall metrics for these tasks were generally near or below 50%, indicating that while models can count, they struggle with precise spatial identification. An exception was Gemini 2.5 Flash in outlier identification, which reached 65.01% Recall.

The study also investigated the impact of chart design on model performance. While chart design was found to be a secondary factor, certain design choices did negatively affect accuracy. Specifically, scatterplots with wide aspect ratios (16:9 and 21:9) or those with randomly colored points led to impaired performance. Conversely, adding opacity to points was found to be beneficial.

Also Read:

Key Takeaways for AI Model Application

Based on these results, the researchers offer several considerations for combining scatterplot images with AI models:

  1. Prioritize few-shot prompting: This strategy consistently yielded better results across all models and tasks.
  2. Exercise caution with localization tasks: Current OpenAI and low-cost Google models are not reliable for precise detection or identification in scatterplots.
  3. Focus on other components before chart design: While chart design is crucial for human understanding, its impact on AI performance is less significant, though avoiding wide aspect ratios and random colors is advisable.

This research marks a crucial step towards better understanding and improving AI models for chart comprehension, particularly for scatterplots. The authors plan to expand the dataset, explore more prompting strategies and tasks, and investigate the trade-offs between cost and performance. They also aim to fine-tune smaller, open models and eventually work on generating accessible alt-text descriptions for chart images. You can find the full research paper here.

Karthik Mehta
Karthik Mehtahttps://blogs.edgentiq.com
Karthik Mehta is a data journalist known for his data-rich, insightful coverage of AI news and developments. Armed with a degree in Data Science from IIT Bombay and years of newsroom experience, Karthik merges storytelling with metrics to surface deeper narratives in AI-related events. His writing cuts through hype, revealing the real-world impact of Generative AI on industries, policy, and society. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -