spot_img
HomeResearch & DevelopmentPISA-Bench: A New Multilingual Benchmark for Evaluating Vision-Language Models

PISA-Bench: A New Multilingual Benchmark for Evaluating Vision-Language Models

TLDR: PISA-Bench is a new, high-quality, multilingual, and multimodal benchmark for evaluating Vision-Language Models (VLMs). Derived from expert-created PISA tests and translated into six languages, it features human-verified examples with images, questions, and answer options. Evaluations show smaller VLMs struggle, and most models experience performance drops on non-English tasks, especially in spatial reasoning, indicating a need for better multilingual multimodal reasoning capabilities in AI.

The field of artificial intelligence has seen remarkable advancements in Vision-Language Models (VLMs), which are systems capable of understanding and reasoning across both images and text. However, evaluating these sophisticated models, especially in diverse linguistic and cultural contexts, has presented a significant challenge. Many existing benchmarks for VLMs often rely on content generated by other large language models, which can limit the quality and diversity of the evaluation data. Furthermore, most of these benchmarks are primarily in English, making it difficult to assess how well VLMs perform in other languages.

To address these critical gaps, researchers Patrick Haller, Fabio Barth, Jonas Golde, Georg Rehm, and Alan Akbik have introduced PISA-Bench. This innovative benchmark is designed to provide a high-quality, human-verified, multilingual, and multimodal metric for evaluating VLMs. PISA-Bench is derived from the internationally recognized PISA tests, which are used to assess student competencies in over eighty countries. This foundation ensures that the benchmark uses real-world, expert-created examples that are designed to be culturally and linguistically unbiased.

Each example within PISA-Bench is meticulously crafted, consisting of human-extracted instructions, questions, answer options, and images. These examples are further enriched with categories that define the type of question, such as spatial and geometric reasoning, quantitative reasoning, graph and pattern analysis, and text and diagram understanding. A key feature of PISA-Bench is its multilingual nature; the English examples have been translated into five additional languages: Spanish, German, Chinese, French, and Italian. This results in a comprehensive, fully parallel corpus covering six languages, allowing for robust cross-lingual evaluation.

The construction of PISA-Bench involved a four-stage pipeline. First, tasks were collected from the original OECD PISA tests. Second, these tasks were decomposed into modular components like instructions, images, questions, and answer options. Third, a rigorous quality assurance step was performed by human annotators to verify and correct the extracted content. Finally, LLM-based translations were generated for the five target languages and then verified by human native speakers. This meticulous process ensures the high quality and reliability of the dataset.

In their evaluation of state-of-the-art vision-language models on PISA-Bench, the researchers made several important findings. They observed that smaller models, particularly those with fewer than 20 billion parameters, struggled significantly, failing to achieve high test scores. Larger and proprietary models performed better, but even these showed substantial performance degradation when evaluated on non-English language splits. This highlights a critical area for improvement in current VLM development.

A particularly challenging area for the models was spatial and geometric reasoning, where error rates were notably high across all languages. This suggests that while VLMs have advanced significantly, they still face difficulties in tasks requiring complex visual interpretation and spatial understanding. The study also included a contamination analysis, which confirmed that the models were not simply memorizing answers from pre-training data, as performance dropped significantly when images were removed. This indicates that PISA-Bench genuinely tests the models’ reasoning capabilities rather than their recall.

Also Read:

By releasing the PISA-Bench dataset and its evaluation framework, the authors provide a valuable resource for advancing research in multilingual multimodal reasoning. This benchmark will enable developers and researchers to better understand the strengths and weaknesses of VLMs, fostering the creation of more capable and globally applicable AI systems. For more detailed information, you can refer to the full research paper here.

Karthik Mehta
Karthik Mehtahttps://blogs.edgentiq.com
Karthik Mehta is a data journalist known for his data-rich, insightful coverage of AI news and developments. Armed with a degree in Data Science from IIT Bombay and years of newsroom experience, Karthik merges storytelling with metrics to surface deeper narratives in AI-related events. His writing cuts through hype, revealing the real-world impact of Generative AI on industries, policy, and society. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -