spot_img
HomeResearch & DevelopmentMCIF: A New Benchmark for Multilingual Multimodal AI Understanding

MCIF: A New Benchmark for Multilingual Multimodal AI Understanding

TLDR: MCIF (Multimodal Crosslingual Instruction Following) is a novel, human-annotated benchmark designed to evaluate multimodal large language models (MLLMs) in complex, real-world scenarios. Based on scientific talks, it assesses instruction-following across three modalities (text, speech, video) and four languages (English, German, Italian, Chinese). MCIF includes 13 diverse tasks, both short and long contexts, and two types of prompts (fixed and varied) to test model robustness. The benchmark highlights current MLLMs’ capabilities and identifies significant performance gaps, particularly in handling long-form content and varied instructions.

Recent advancements in large language models (LLMs) have paved the way for Multimodal LLMs (MLLMs), which can process and understand information from various sources like text, speech, and vision. As these MLLMs become more sophisticated and move towards general-purpose instruction-following, there’s a growing need to evaluate their capabilities, especially in handling multiple languages and different types of content, both short and long.

However, existing evaluation tools often fall short. Many are limited to English, focus on only one type of data at a time, use only short pieces of information, or lack human-created annotations, which makes it difficult to truly assess how well these models perform across different languages, data types, and complex tasks.

To address these critical gaps, researchers have introduced MCIF, which stands for Multimodal Crosslingual Instruction Following. This is the first benchmark of its kind that is human-annotated and built using scientific talks. It’s specifically designed to test how well MLLMs can follow instructions in settings that involve multiple languages and multiple types of data, whether the input is short or long.

What is MCIF?

MCIF is a comprehensive benchmark that covers three main types of data: speech, video, and text. It also spans four diverse languages: English, German, Italian, and Chinese. This broad coverage allows for a thorough evaluation of MLLMs’ ability to understand instructions across different languages and combine them with information from various data sources.

The benchmark is built around 13 distinct tasks, categorized into four main areas: recognition (like transcribing speech), translation (converting content from one language to another), question answering (finding answers to questions based on provided content), and summarization (creating shorter versions of content). It includes both short-form and long-form content, allowing for evaluation of how models handle different context lengths.

A key strength of MCIF is its manual curation and human annotations. Professional linguists created high-quality transcripts in English. For summarization, abstracts from the associated scientific papers were used. For question answering, expert annotators, who are also authors of the paper, created over 200 unique question-and-answer pairs. These questions were designed to test different aspects of understanding, including general knowledge, information from transcripts, and information from abstracts. Each question was also labeled based on whether the answer could be found in audio, video, both, or neither.

To ensure crosslingual capabilities, all English textual data, including transcripts, summaries, and Q&A pairs, were professionally translated into Italian, German, and Chinese. This rigorous process ensures high data quality and consistency.

Testing Instruction Following

MCIF also introduces two sets of prompts to evaluate models: MCIFfix and MCIFmix. MCIFfix uses a consistent, fixed prompt for each task area, while MCIFmix includes a variety of different ways to phrase the same instruction. This allows researchers to measure how robust models are to variations in how instructions are given, which is crucial for real-world applications where users might phrase requests differently.

Evaluating Performance

The researchers evaluated 21 state-of-the-art models, including traditional Large Language Models (LLMs), SpeechLLMs (specialized in speech), VideoLLMs (specialized in video), and Multimodal LLMs (MLLMs) that handle all data types. They used standard metrics like Word Error Rate (WER) for recognition, COMET for translation, and BERTScore for question answering and summarization.

Initial results show that while SpeechLLMs perform well on short audio recognition, their performance drops significantly with longer audio. MLLMs like Ola, however, showed impressive performance on long-form recognition. In translation, LLMs generally excelled, but many SpeechLLMs and MLLMs struggled with long-form content, often failing to translate the entire context. For question answering, MLLMs like Ola and Qwen2.5-Omni were highly competitive, especially with long-form inputs, while LLMs showed limitations in crosslingual information extraction.

A notable finding was the consistent drop in performance across tasks when models were tested with the varied prompts of MCIFmix compared to the fixed prompts of MCIFfix. This highlights that many current models lack robustness to different ways of phrasing instructions, a key area for future improvement.

Also Read:

Conclusion

MCIF represents a significant step forward in evaluating the capabilities of multimodal and multilingual instruction-following systems. By providing a human-annotated benchmark with diverse modalities, languages, tasks, and context lengths, it offers a comprehensive framework for understanding the strengths and weaknesses of current MLLMs. The benchmark is openly available under a CC-BY 4.0 license to encourage further research and development in this exciting field. You can find the full research paper here: MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -