spot_img
HomeResearch & DevelopmentAddressing Semantic Confusion in AI's Long Video Comprehension

Addressing Semantic Confusion in AI’s Long Video Comprehension

TLDR: The paper introduces ELV-Halluc, the first benchmark for Semantic Aggregation Hallucination (SAH) in long video understanding. SAH occurs when video AI models correctly perceive frame-level details but misattribute them across different events in a long video. The research confirms SAH’s existence, shows it increases with semantic complexity and rapid semantic changes, and proposes mitigation strategies like improved positional encoding and Direct Preference Optimization (DPO), achieving a significant reduction in SAH.

Video multimodal large language models (Video-MLLMs) have made impressive strides in understanding video content. However, a significant challenge persists: hallucination. This is when these models generate information that is inconsistent with or entirely unrelated to the video input. While previous benchmarks for video hallucination primarily focused on short videos, attributing errors to factors like strong language biases or missing frames, a new type of hallucination has been identified, particularly critical in longer videos.

This new phenomenon is called Semantic Aggregation Hallucination (SAH). SAH occurs when models correctly identify individual elements or actions at the frame level but then incorrectly combine or attribute these semantics across different events within a longer video. Imagine a news broadcast where a model correctly identifies a ‘Starbucks cup’ in one segment and ‘papers’ in another, but then incorrectly states that the host is holding a ‘Starbucks cup’ while explaining the news, even though they were holding papers at that specific moment. The individual elements are seen, but their temporal context and aggregation are flawed.

Given that long videos inherently possess increased semantic complexity due to multiple distinct events, SAH becomes a more pronounced issue. To address this, researchers have introduced ELV-Halluc, the first benchmark specifically designed to investigate long-video hallucination and, more precisely, SAH. This benchmark allows for a systematic study of how and why these semantic aggregation errors occur.

The ELV-Halluc benchmark is built upon ‘event-by-event’ videos, which are long videos composed of several clearly separated events sharing a common theme, like a news report with multiple distinct news items. This structure helps in isolating semantic units and quantifying semantic complexity based on the number of events. The dataset includes 348 high-quality videos with human-refined captions, from which 200 videos are selected for the benchmark.

To evaluate SAH, ELV-Halluc uses an adversarial question-answer pair design. For each ground-truth caption, two types of hallucinated captions are created: ‘in-video’ and ‘out-video’. In-video modifications replace an element (like an object or action) with something that *does* appear in another part of the same video. Out-video modifications replace an element with something fabricated that *does not* appear anywhere in the video. By comparing a model’s accuracy on these two types of questions, researchers can quantify the contribution of SAH. A model that performs well on out-video hallucinations (rejecting fabricated content) but poorly on in-video ones (being misled by semantically plausible but temporally misattributed content) indicates a susceptibility to SAH.

Experiments conducted on 14 open-source and 2 closed-source Video-MLLMs using ELV-Halluc confirmed the widespread existence of SAH. Key findings include:

Also Read:

Key Findings on SAH

  • SAH positively correlates with semantic complexity; models are more prone to SAH as the number of events in a video increases.
  • SAH occurs more frequently on rapidly changing semantics, such as visual details, followed by actions, then objects, and least on declarative content. This suggests that models struggle more with fine-grained, temporally dynamic information.
  • Increasing the number of sampled frames generally improves overall hallucination accuracy but can also increase the SAH ratio, as richer information also introduces more opportunities for semantic mismatches.
  • Larger models tend to have higher overall hallucination accuracy, but model size does not consistently correlate with a reduction in SAH ratio.

The research also explores potential strategies to mitigate SAH. One approach involves strengthening the positional encoding mechanism, which helps models bind semantic relationships more effectively across temporal segments. Experiments showed that advanced positional encoding strategies like VideoRoPE can contribute to lowering the SAH ratio.

Another promising mitigation strategy is Direct Preference Optimization (DPO). By training models with specific positive (ground truth) and negative (in-video hallucinated) response pairs, the model’s preference for hallucinated semantics can be suppressed. Applying DPO with in-video hallucinated pairs significantly reduced the SAH ratio by 27.7% and also slightly improved general video understanding performance on other benchmarks. This indicates that explicitly teaching the model to distinguish correct event semantics from misaggregated ones is highly effective.

While ELV-Halluc represents a crucial step forward in understanding and addressing hallucinations in long video AI, the authors acknowledge limitations such as potential biases from semi-automated caption generation and the controlled nature of event-based videos compared to real-world long videos. Nevertheless, this benchmark and the proposed mitigation strategies lay a solid foundation for developing more reliable long-video understanding models. You can read the full research paper here.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -