spot_img
HomeResearch & DevelopmentDixit: A New Frontier for Evaluating Multimodal AI Capabilities

Dixit: A New Frontier for Evaluating Multimodal AI Capabilities

TLDR: A research paper proposes using the fantasy card game Dixit as a novel, holistic, and engaging benchmark for evaluating Multimodal Large Language Models (MLMs). It finds that MLMs perform significantly better than random players, their Dixit rankings correlate perfectly with established benchmarks, and even a weaker MLM can compete with human players. The study also reveals that MLMs tend to use literal captioning strategies compared to humans’ more abstract and external-knowledge-based approaches, highlighting areas for future AI improvement in common-sense reasoning and nuanced understanding.

A new research paper explores an innovative approach to evaluating Multimodal Large Language Models (MLMs) by using the popular fantasy card game, Dixit. Titled “Can They Dixit? Yes they Can! Dixit as a Playground for Multimodal Language Model Capabilities,” the paper highlights the limitations of current MLM evaluation methods and proposes game-based assessments as a more holistic and engaging alternative.

Current methods for assessing MLMs often fall into two categories: static, individual benchmarks that test capabilities in isolation, or pairwise comparisons relying on human or model judgments. The authors, Nishant Balepur, Dang Nguyen, and Dayeon Ki from the University of Maryland, point out that these methods can be subjective, expensive, and allow models to exploit superficial shortcuts like verbosity to inflate their performance.

The Dixit Advantage

To overcome these challenges, the researchers advocate for game-based evaluations. Games, by their nature, require multiple abilities to win, are inherently competitive, and are governed by fixed, objective rules. This framework makes evaluations more engaging and robust. Dixit was chosen as the specific game for this evaluation suite because it effectively assesses several diverse MLM capabilities simultaneously.

In Dixit, one player, the ‘storyteller,’ selects a card from their hand and generates a caption for it. The goal is to create a caption that is relevant enough for some, but not all, other players to guess the correct card. Other players then choose a card from their own hand that they believe best matches the caption. All selected cards are pooled, and players vote on which card they think belongs to the storyteller. Points are awarded based on successful guesses and misdirections.

This game structure demands a range of MLM skills: generating creative captions (testing creativity, image comprehension, and calibration), and identifying cards that align with the storyteller’s caption while also considering what other players might choose (testing accuracy, image comprehension, and theory-of-mind). The paper argues that Dixit provides a unified and nuanced evaluation suite for these complex abilities.

Experimental Setup and Key Findings

The researchers built their own evaluation interface for Dixit and had five prominent MLMs compete: GPT-4o, Claude-3.5 Sonnet, Intern-VL2, Qwen2VL, and Molmo. They also included a ‘Random’ player for baseline comparison. The images used were from an open-source version of Dixit to avoid copyright issues.

The experiments yielded several significant findings. Firstly, all MLMs significantly outperformed the Random player, demonstrating that their acquired skills in image comprehension, reasoning, and classification are transferable to an out-of-domain task like Dixit. Secondly, the rankings of MLMs based on their Dixit performance perfectly correlated with their rankings on popular leaderboards such as ChatbotArena and the Open VLM Leaderboard. This suggests that game-based evaluations like Dixit can comprehensively test MLM capabilities within a single, unified task.

In a fascinating human-versus-MLM experiment, the authors played Dixit against GPT-4o Mini (a weaker MLM). While humans generally had the upper hand, GPT-4o Mini managed to beat one human player on average points. This indicates that MLMs possess strong Dixit capabilities that could potentially rival or even surpass human abilities, especially with more capable models.

Also Read:

MLM Strategies and Areas for Improvement

A detailed analysis of MLM strategies revealed interesting differences. MLMs tended to generate longer, more literal captions (e.g., “Child reaching for the moon”), directly describing what was in the picture. In contrast, human players often produced shorter, more abstract captions that referenced external knowledge or pop culture (e.g., “Rapunzel”). This suggests that future MLMs could improve their Dixit performance by enhancing their common-sense reasoning capabilities and ability to generate more ambiguous, externally-referenced captions.

The study also looked into MLM reasoning errors when selecting cards. While GPT-4o showed high accuracy, other models exhibited issues like implausible rationales, hallucinations (referencing things not in the card), or a complete lack of reasoning. This highlights Dixit’s potential as a testbed for studying MLM hallucinations and reasoning gaps.

The paper concludes by advocating for game-based frameworks like Dixit as robust, challenging, and engaging testbeds for evaluating MLM capabilities holistically. It encourages future work to explore training strategies beyond basic prompting to build stronger Dixit models, potentially through self-play, and to design models that not only win but also make the game enjoyable for human players. For more details, you can read the full research paper here.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -