spot_img
HomeResearch & DevelopmentEgoNight: Advancing Egocentric AI in Low-Light Conditions

EgoNight: Advancing Egocentric AI in Low-Light Conditions

TLDR: EgoNight is the first comprehensive benchmark for egocentric vision understanding in low-light conditions. It features day-night aligned videos from synthetic and real-world sources, a visual question answering (VQA) dataset with 3,658 human-verified pairs across 12 types, and auxiliary tasks like day-night correspondence retrieval and depth estimation. Experiments show that current multimodal large language models (MLLMs) perform significantly worse at night, highlighting the need for more robust AI systems in challenging illumination.

Most of our daily lives, and consequently, the data collected for artificial intelligence systems, occur during the day or in well-lit environments. However, real-world applications of AI, especially in areas like robotics, autonomous driving, and personal assistance, often require systems to function effectively in low-light or nighttime conditions. This significant gap in current research is precisely what the new EgoNight benchmark aims to address.

The research paper, titled “EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark,” introduces the first comprehensive benchmark specifically designed for egocentric vision understanding at night. Egocentric vision refers to a first-person perspective, much like what a person wearing a camera would see. This field is crucial for developing intelligent assistants that can perceive and interact with the world as humans do.

What is EgoNight?

EgoNight is a groundbreaking benchmark that integrates diverse video sources, including synthetic environments and real-world indoor and outdoor scenes. A key innovation is the inclusion of day-night aligned videos. This means that for many scenarios, the same scene and actions are captured under both daytime and nighttime conditions, allowing for a direct and rigorous comparison of how AI models perform under varying illumination. This alignment is vital for understanding the specific challenges posed by low light.

The dataset is composed of three main sources: EgoNight-Synthetic, generated using Blender for perfect day-night alignment in indoor scenes; EgoNight-Sofia, real-world recordings from Sofia, Bulgaria, capturing diverse daily activities with reasonable day-night alignment; and EgoNight-Oxford, which incorporates existing nighttime videos from the Oxford Day-and-Night dataset to add scale and diversity, particularly for urban outdoor scenes.

Core Tasks and Challenges

The primary focus of EgoNight is egocentric Visual Question Answering (VQA). This involves asking questions about a video and expecting an accurate answer, reflecting a high-level understanding of the visual content. EgoNight-VQA features 3,658 human-verified question-answer pairs across 12 diverse types of questions. These range from well-studied tasks like object recognition, text recognition, and spatial reasoning to newly proposed and more challenging categories such as scene sequence understanding, navigation, lighting recognition, and even non-common-sense reasoning (e.g., identifying physically implausible scenarios in synthetic data).

The questions are categorized into “paired” types, where the same question applies to both day and night videos, and “unpaired” types, which are only meaningful or practical to ask in nighttime conditions (like lighting recognition or dynamic detection). This meticulous design allows researchers to pinpoint exactly where models struggle when light diminishes.

Beyond VQA, EgoNight also introduces two auxiliary tasks: day-night correspondence retrieval, which tests a model’s ability to match visual content across different lighting conditions, and egocentric depth estimation at night, a critical capability for embodied AI systems that need to understand their 3D environment in the dark.

Also Read:

Key Findings and Future Directions

Extensive evaluations of state-of-the-art multimodal large language models (MLLMs), including commercial systems like GPT-4.1 and Gemini 2.5 Pro, revealed a consistent and substantial drop in performance when transferring from day to night conditions. This highlights that current MLLMs are not robust to illumination changes and struggle significantly with reasoning in low-light scenarios. Perception-oriented tasks, such as object and text recognition, showed larger performance drops, indicating their higher sensitivity to illumination, while reasoning tasks, though harder overall, were relatively less affected.

The newly proposed QA types, especially those related to lighting recognition, dynamic detection, navigation, and non-common-sense reasoning, proved to be particularly challenging for existing MLLMs, suggesting that these models generalize poorly to novel tasks in low-light settings.

The EgoNight benchmark, developed by authors including Deheng Zhang, Yuqian Fu, and Danda Pani Paudel, provides a crucial foundation for advancing research in egocentric vision. It motivates the development of more robust AI models that can generalize across different illumination domains, ultimately leading to more reliable and capable AI assistants for real-world applications. All data and code for EgoNight will be made publicly available upon acceptance of the paper, fostering collaborative progress in this vital area. You can find more details in the full research paper available at arXiv:2510.06218.

Karthik Mehta
Karthik Mehtahttps://blogs.edgentiq.com
Karthik Mehta is a data journalist known for his data-rich, insightful coverage of AI news and developments. Armed with a degree in Data Science from IIT Bombay and years of newsroom experience, Karthik merges storytelling with metrics to surface deeper narratives in AI-related events. His writing cuts through hype, revealing the real-world impact of Generative AI on industries, policy, and society. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -