TLDR: A new benchmark dataset, PixelHumor, consisting of 2,800 annotated multi-panel comics, evaluates Large Multimodal Models’ (LMMs) ability to understand visual and narrative humor. Experiments show significant limitations in LMMs, particularly in panel sequencing, humor interpretation, and classifying subtle humor styles, indicating they struggle with complex multimodal reasoning compared to humans.
Understanding humor is a complex human trait, involving a blend of social intelligence, abstract thinking, and contextual reasoning. While Large Language Models (LLMs) have shown impressive capabilities in text-based tasks, grasping humor, especially visual humor, remains a significant hurdle. This is where Large Multimodal Models (LMMs), which can process both text and images, come into play. However, even with these advanced models, there has been a lack of systematic evaluation for their ability to interpret humor in multi-panel comics, which often rely on a combination of visual and textual cues to deliver a punchline.
To address this challenge, researchers from the Singapore University of Technology and Design have introduced a new benchmark dataset called PixelHumor. This dataset comprises 2,800 annotated multi-panel comics, specifically designed to test how well LMMs can interpret multimodal humor and understand narrative sequences. The findings from their experiments with state-of-the-art LMMs highlight substantial gaps in current models’ abilities, indicating that even top models struggle significantly with tasks like panel sequencing and humor interpretation.
Introducing PixelHumor: A New Benchmark for Comic Understanding
PixelHumor is a comprehensive dataset built from seven diverse online comic sources, including popular titles like Cyanide and Happiness, Peanuts, Garfield, and XKCD. These comics were chosen to represent a broad spectrum of humor styles, such as Comparison, Personification, Exaggeration, Pun, Sarcasm, Silliness, Surprise, and Dark humor. Each comic in the dataset has been meticulously annotated for various humor-related tasks, providing a rich resource for evaluating AI systems. The dataset itself is made available through URLs linking to the original sources, respecting copyright and creator hosting.
Evaluating LMMs: The Four Core Tasks
The researchers defined four core tasks to thoroughly evaluate LMMs’ humor comprehension:
Humor Identification: This task assesses whether a model can detect humor, identify contributing factors like sound effects, pinpoint the most critical panel for humor, and determine if text or visuals are more important. While models performed well in simply detecting the presence of humor, they struggled with attributing its source or identifying the key panel.
Humor Classification: Models were asked to categorize comics into one or more of the eight predefined humor styles. GPT-4o performed best, but many open-source models showed a bias towards assigning only a single humor style, often misclassifying nuanced types like Sarcasm and Dark humor.
Humor Interpretation: This open-ended task required models to generate natural language explanations for why a comic is humorous. GPT-4o again led the pack, but even top models showed volatility in longer narratives, indicating difficulties in maintaining multimodal coherence over extended sequences. Human evaluators consistently preferred human-written explanations, highlighting a significant gap in AI’s ability to “think” about humor in human-like terms.
Sequence Recognition: Comics rely on a specific panel order to build context and deliver punchlines. This task evaluated models’ ability to reconstruct the correct visual and textual sequence. Gemini-1.5-Pro and GPT-4o achieved the highest accuracy, but all models struggled, often defaulting to conventional reading orders rather than understanding the narrative flow, especially as the number of panels increased.
Also Read:
- AI’s Physical Intuition: Why Large Multimodal Models Struggle to Learn New Physics from Visual Examples
- Fables Expose Flaws in Large Language Models’ Moral Understanding
Key Findings and Future Directions
The experiments revealed that while LMMs are good at recognizing the presence of humor, they fall short in deeper understanding, such as narrative sequencing, subtle humor classification, and attributing humor to specific modalities. Performance generally degraded as comic complexity and panel count increased, suggesting that current LMMs often rely on surface-level cues rather than true multimodal contextual and sequential reasoning.
This research underscores the need for future AI systems to move beyond simple pattern recognition. To genuinely comprehend humor, models need to understand the causal, temporal, and multimodal relationships that form comic narratives. This could involve developing hierarchical modeling of narrative structures and improved cross-modal fusion mechanisms to better integrate visual and textual cues. The PixelHumor benchmark, detailed further in the research paper, aims to drive these advancements, paving the way for AI systems that can engage more socially intelligently with complex human communication like humor.


