TLDR: ToM-SSI is a novel, multimodal benchmark designed to evaluate AI’s Theory of Mind in complex, situated social interactions involving up to four agents. It moves beyond traditional, limited tests by incorporating group dynamics and mixed cooperative-obstructive scenarios. Evaluations reveal that current AI models significantly underperform compared to humans, struggling with inferring perceptions, beliefs, and intentions, and often failing to effectively integrate visual information. The benchmark highlights critical areas for future research in developing more socially intelligent AI.
The ability to understand and attribute mental states like beliefs, intentions, desires, and knowledge to oneself and others, known as Theory of Mind (ToM), is crucial for effective human social interaction, empathy, and communication. As artificial intelligence (AI) models become more sophisticated, evaluating their ToM capabilities is essential for developing truly intelligent and socially aware systems.
However, existing benchmarks for assessing AI’s Theory of Mind often fall short. Many are based on simplified scenarios, like variations of the classic Sally-Anne test, which offer a very limited view of social cognition. Furthermore, these benchmarks are frequently text-only or involve interactions between just two agents, neglecting the rich complexity and spatial dynamics inherent in real-world human social interactions.
To address these significant gaps, researchers Matteo Bortoletto, Constantin Ruhdorfer, and Andreas Bulling from the University of Stuttgart, Germany, have introduced a novel benchmark called ToM-SSI: Theory of Mind in Situated Social Interactions. This new evaluation framework is specifically designed to test AI models in environments that are rich with social interactions and spatial elements.
ToM-SSI stands out by being inherently multimodal, combining visual and textual information. Unlike previous benchmarks, it supports group interactions involving up to four agents who communicate and move within a grid-world environment. This unique design allows for the study of complex scenarios, including mixed cooperative-obstructive settings, where agents might be collaborative with some and obstructive towards others. It also enables the parallel reasoning about multiple agents’ mental states, capturing a much broader spectrum of social cognition than ever before.
The benchmark comprises five distinct tasks: Cooperative Movement – Single Communication (CMSC), Cooperative Movement – Concurrent Communication (CMCC), Probabilistic Cooperative Communication (PCC), Obstructive Communication (OC), and Mixed Cooperative-Obstructive Communication (MC). Each task is paired with three types of questions designed to probe different aspects of an agent’s mental state: percepts (what an agent observes), beliefs (an agent’s internal representation of the world based on percepts and prior knowledge), and intentions (the actions an agent commits to based on their beliefs and desires). This causal structure—Percept leading to Belief, and Belief leading to Intention—is critical for a comprehensive evaluation of ToM.
The evaluations conducted using ToM-SSI revealed several important insights into the current state of AI’s social intelligence. A primary finding is that state-of-the-art large foundation models perform significantly worse than humans across all tasks. For some tasks, their performance even falls below that of smaller models, highlighting a substantial gap in their reasoning abilities.
Models particularly struggle with the sequential steps required for robust ToM reasoning. While they might perform reasonably well in inferring an agent’s percepts, their accuracy drops considerably when asked to determine beliefs based on those percepts, and declines even further when inferring intentions from beliefs. This suggests a fundamental challenge in consistently integrating and processing information to build a coherent understanding of an agent’s mental state.
An in-depth error analysis further pinpointed specific limitations. Models showed difficulties in accurately modeling agent perception, handling multi-agent communication, and navigating mixed social interactions. For instance, they sometimes failed to account for agents observing each other’s initial knowledge or overlooked the implications of information being passed between agents in a group setting.
Perhaps one of the most surprising findings was that most Vision-Language Models (VLMs) did not benefit from the inclusion of image inputs. In some cases, models like GPT-4o even performed better on text-only versions of the tasks. This indicates a critical disparity in how these models leverage multimodal information, suggesting a gap in their ability to effectively integrate visual cues with textual descriptions to perform ToM tasks.
While ToM-SSI utilizes a synthetic grid-world environment, which is simpler than the real world, this design choice allows researchers to focus on core ToM abilities without the confounding factors of complex common-sense reasoning or hallucinations often encountered in more realistic simulations. The benchmark’s findings underscore critical areas for future research, particularly in improving AI’s ability to model agent perception, understand multi-agent communication, and handle the nuances of mixed social interactions.
Also Read:
- ProToM: An AI That Understands and Encourages Helpful Behavior Among Independent Agents
- Evaluating AI: Bridging the Gap Between Benchmarks and Human Understanding
For more detailed information, you can read the full research paper here.


