spot_img
HomeResearch & DevelopmentUnveiling Fairness in Multilingual AI: A New Benchmark for...

Unveiling Fairness in Multilingual AI: A New Benchmark for Visual Question Answering

TLDR: LinguaMark is a new benchmark evaluating Large Multimodal Models (LMMs) for fairness, relevancy, and faithfulness in multilingual Visual Question Answering (VQA). It uses 6,875 image-text pairs across 11 languages and five social attributes. The study found that closed-source models generally outperform open-source ones, though open-source models like Qwen2.5 show strong generalization. It highlights persistent biases, especially in gender-related prompts and low-resource languages like Tamil and Urdu, emphasizing the need for culturally aware AI evaluation.

Large Multimodal Models (LMMs) have rapidly advanced, enabling impressive capabilities in understanding both images and text. However, a recent research paper titled “LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation” by Ananya Raval, Aravind Narayanan, Vahid Reza Khazaie, and Shaina Raza, sheds light on a critical limitation: these models often struggle with linguistic diversity, leading to biased and unfair outputs across different languages.

The core issue is that most LMMs are predominantly trained on vast amounts of data from high-resource languages like English, Mandarin, and Spanish. This focus leaves significant gaps in their performance when dealing with less common, or ‘low-resource,’ languages. For instance, while an LMM might excel at answering questions about an image in English, its accuracy can drop considerably when asked the same question in a language with fewer available training resources. This disparity is particularly evident in visually grounded tasks such as image captioning or Visual Question Answering (VQA).

Existing benchmarks for evaluating LMMs, such as MM-Vet and MMBench, primarily focus on overall accuracy, especially in high-resource languages. They often overlook crucial aspects like linguistic fairness, cultural bias, and how faithfully a model’s answer aligns with visual evidence across a diverse set of languages. This creates a significant blind spot in understanding how LMMs perform in real-world, socially sensitive multilingual scenarios, particularly concerning bias and stereotyping, factual consistency with images, and relevance to the given prompt.

Introducing LinguaMark: A New Standard for Multilingual Evaluation

To address this critical gap, the researchers introduce LinguaMark, a groundbreaking multilingual benchmark designed as an open-ended Visual Question Answering (VQA) task. LinguaMark provides two key components: a carefully curated multilingual test set and a standardized evaluation framework for both open-source and closed-source LMMs.

The dataset is impressive, comprising 6,875 unique image-text pairs. These pairs are adapted from prior work and meticulously translated into 11 languages: English, Bengali, French, Korean, Mandarin, Persian, Portuguese, Punjabi, Spanish, Tamil, and Urdu. Crucially, all annotations and translations are human-verified to ensure linguistic and cultural fidelity. The VQA pairs are also categorized under five demographic social attributes: age, gender, race, occupation, and sports, allowing for a detailed analysis of bias.

LinguaMark evaluates models using three key metrics:

  • Bias: Measures the degree of social bias in model output across protected attributes. Lower values indicate reduced biased behavior.

  • Answer Relevancy: Assesses how factually correct the model is in identifying the image and producing an accurate natural language output.

  • Faithfulness: Determines how well the model’s answer aligns with the ground truth answer in its respective language, indicating multilingual fluency.

Key Findings: Closed-Source Models Lead, Open-Source Show Promise

The comprehensive evaluation included leading LMMs, both closed-source (like GPT-4o and Gemini 2.5 Flash) and open-source (such as Qwen2.5-Vision-Instruct, Aya-Vision-8B, Gemma3-12B-it, Llama-3.2-11B-Vision-Instruct, and Phi-4-multimodal-instruct). The findings reveal several important insights:

  • Overall Performance: Closed-source models, particularly Gemini 2.5 Flash and GPT-4o, consistently outperformed their open-source counterparts in terms of answer relevancy and faithfulness. Gemini 2.5 Flash achieved the highest scores in both, demonstrating strong generalization across multiple languages and modalities.

  • Bias Levels: While bias levels were relatively consistent across models, GPT-4o showed the lowest bias overall. However, the study found that the ‘Gender’ social attribute consistently exhibited the highest bias values across all models, followed by ‘Age’ and ‘Occupation’. ‘Sports’ and ‘Ethnicity’ had the lowest bias.

  • Language Disparities: English, being the most dominant language in training corpora, naturally performed best in bias and answer relevancy. However, low-resource languages like Tamil and Urdu showed some of the highest bias scores and lowest answer relevancy and faithfulness, highlighting the persistent challenges in these linguistic contexts.

  • Open-Source Generalization: Notably, the open-source model Qwen2.5 demonstrated strong generalization capabilities, achieving minimal bias scores even in languages it wasn’t explicitly trained on, such as Bengali and Spanish.

The research also includes qualitative examples, showing how models respond to questions in different languages, even those they might not have been explicitly trained on, and how they interpret culturally sensitive visual information.

Limitations and Future Directions

The researchers acknowledge several limitations. LinguaMark currently focuses on a relatively small set of languages, and the dataset primarily uses images from news articles and social media. Future work aims to expand the dataset to include more diverse and critical categories like surveillance and medical contexts. They also plan to incorporate larger open-source LMMs (exceeding 14B parameters) for a more direct comparison with large-scale closed-source models and to broaden the evaluation to include additional tasks beyond open-ended VQA, such as close-ended VQA and sentiment analysis.

Also Read:

Conclusion

LinguaMark represents a significant step forward in evaluating the fairness, relevancy, and faithfulness of LMMs in multilingual settings. While closed-source models currently lead in overall accuracy and alignment, open-source models like Qwen2.5 show promising generalization, especially in low-resource languages. The persistent disparities across languages and social categories, particularly in gender-based prompts and underrepresented languages, underscore the urgent need for culturally aware evaluation and greater model transparency. By releasing their benchmark and evaluation code, the authors hope to encourage broader adoption of fairness-aware evaluation in multimodal systems and inspire future improvements in both open and proprietary models. You can find the full research paper here.

Karthik Mehta
Karthik Mehtahttps://blogs.edgentiq.com
Karthik Mehta is a data journalist known for his data-rich, insightful coverage of AI news and developments. Armed with a degree in Data Science from IIT Bombay and years of newsroom experience, Karthik merges storytelling with metrics to surface deeper narratives in AI-related events. His writing cuts through hype, revealing the real-world impact of Generative AI on industries, policy, and society. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -