spot_img
HomeResearch & DevelopmentBridging the Viewpoint Gap: Enhancing Video Language Models for...

Bridging the Viewpoint Gap: Enhancing Video Language Models for Consistent Temporal Understanding

TLDR: A new benchmark, EgoExo-Con, reveals that current Video-LLMs struggle with consistent temporal understanding across different camera viewpoints (egocentric and exocentric). To address this, researchers propose View-GRPO, a reinforcement learning framework that improves cross-view consistency by encouraging viewpoint-specific reasoning while aligning overall comprehension, significantly outperforming traditional fine-tuning methods.

Recent advancements in video large language models (Video-LLMs) have showcased impressive abilities in tasks like question answering and temporal grounding. However, a significant challenge remains: can these models consistently understand events when they are captured from different camera viewpoints? This question is crucial because the same event can look drastically different from a first-person (egocentric) view compared to a third-person (exocentric) view, yet the underlying actions and their timing are identical.

Introducing EgoExo-Con: A New Benchmark

To thoroughly investigate this, researchers have introduced EgoExo-Con (Consistency), a novel benchmark designed to evaluate how well Video-LLMs maintain consistent temporal understanding across varying perspectives. EgoExo-Con features 491 comprehensively synchronized egocentric and exocentric video pairs, accompanied by human-refined natural language queries. The benchmark focuses on two key temporal understanding tasks: Temporal Verification, which is a binary question asking if an event occurs within a specific video moment, and Temporal Grounding, which requires identifying the precise start and end timestamps of an event.

What makes EgoExo-Con unique is its emphasis on evaluating not just the correctness of predictions but also their consistency across viewpoints. The data for this benchmark is sourced from diverse datasets like CharadesEgo, LEMMA, and Ego-Exo4D, covering a wide range of daily human-object interactions and skilled tasks. To ensure reliable evaluation, the queries undergo a multi-stage refinement process, including conversion of raw labels into natural language, enrichment by large models like GPT-4o, and rigorous human validation to eliminate ambiguities and ensure accuracy across both views. Misaligned queries are also generated to serve as negative samples for temporal verification, balancing the task.

Key Challenges and Model Performance

The analysis using EgoExo-Con revealed critical limitations in existing Video-LLMs. A major finding is that models often struggle to maintain consistency across viewpoints, with their cross-view performance being significantly worse than their single-view capabilities. This suggests that models might be relying on view-specific biases rather than developing a robust, view-invariant understanding of temporal events.

Furthermore, simply fine-tuning models with synchronized videos from both viewpoints did not reliably improve consistency. In some cases, this ‘naive’ multi-view fine-tuning even led to underperformance compared to models trained on a single view, indicating that merging viewpoints without careful consideration can introduce conflicting information and undermine temporal signals. While advanced closed-source models like GPT-5 and Gemini-2.5 Flash generally outperformed open-source counterparts, a substantial gap (37-47%) in cross-view consistency still exists when compared to human performance, highlighting the difficulty of the benchmark and ample room for improvement. Interestingly, the study also found that effective temporal reasoning and modeling are more crucial than merely increasing the number of input frames.

View-GRPO: A Novel Approach for Consistent Understanding

To address these limitations, the researchers propose View-GRPO, a novel reinforcement learning (RL) framework. View-GRPO is designed to strengthen view-specific temporal reasoning while simultaneously encouraging consistent comprehension across different viewpoints. It builds upon Group Relative Policy Optimization (GRPO), which uses relative rewards within a group of responses to stabilize optimization.

The core of View-GRPO involves training models with carefully curated data called View30K, which includes step-by-step temporal reasoning chains generated by GPT-5 for both egocentric and exocentric views. These reasoning chains are designed to be viewpoint-specific but converge to consistent final answers. The framework uses a comprehensive reward system that includes a format reward (for structured output), an accuracy reward (based on temporal Intersection-over-Union for grounding or binary correctness for verification), and a reasoning reward. The reasoning reward, provided by an LLM-judge, measures how closely the model’s generated reasoning aligns with the target reasoning chains, ensuring high-quality, faithful explanations.

Also Read:

The Impact of View-GRPO

Experiments demonstrate that View-GRPO consistently outperforms naive supervised fine-tuning (SFT) and a basic GRPO implementation, especially in improving cross-view consistency. The reasoning reward component is identified as playing a central role, as it helps models reduce their reliance on view-specific biases and instead learn shared temporal abstractions by encouraging detailed, viewpoint-tailored explanations that lead to consistent conclusions. This approach marks a significant step towards achieving more robust and consistent video understanding, independent of the camera’s perspective.

For more in-depth information, you can read the full research paper here.

Nikhil Patel
Nikhil Patelhttps://blogs.edgentiq.com
Nikhil Patel is a tech analyst and AI news reporter who brings a practitioner's perspective to every article. With prior experience working at an AI startup, he decodes the business mechanics behind product innovations, funding trends, and partnerships in the GenAI space. Nikhil's insights are sharp, forward-looking, and trusted by insiders and newcomers alike. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -