TLDR: The 9th AI City Challenge, hosted at ICCV 2025, significantly advanced real-world AI applications in transportation, industrial automation, and public safety. Featuring four tracks—multi-camera 3D tracking, traffic safety video analysis, warehouse spatial reasoning, and fisheye road object detection—the challenge saw a 17% increase in participation and over 30,000 dataset downloads. It emphasized scalable, efficient, and actionable AI solutions, with top teams leveraging synthetic data, large vision-language models, and edge-optimized designs to set new benchmarks in diverse computer vision tasks.
The 9th AI City Challenge, hosted at ICCV 2025, marked a significant stride in applying computer vision and artificial intelligence to real-world scenarios in transportation, industrial automation, and public safety. This annual event brings together researchers and developers to tackle complex challenges, and the 2025 edition saw remarkable growth, with a 17% increase in participation, attracting 245 teams from 15 countries.
A notable development this year was the public release of several challenge datasets on Hugging Face, leading to over 30,000 downloads. This move significantly broadened access to the challenge, fostering wider community engagement and accelerating research in these critical domains. The challenge emphasized developing scalable and actionable AI solutions, pushing the boundaries of multimodal scene understanding, 3D spatial reasoning, and real-time perception.
Four Key Tracks Driving Innovation
The 2025 Challenge featured four distinct tracks, each addressing a crucial frontier in applied AI:
Multi-Camera 3D Perception (Track 1): This track focused on advancing multi-camera tracking into the 3D domain. Participants were tasked with tracking diverse object types, including people, service robots, and forklifts, across complex indoor layouts using a large-scale synthetic dataset generated via NVIDIA Omniverse. The challenge provided detailed 3D bounding boxes and camera calibration data, with evaluation based on the 3D Higher Order Tracking Accuracy (HOTA) metric.
Traffic Safety Description and Analysis (Track 2): This track challenged teams to perform detailed video captioning and video question answering (VQA) on staged traffic scenarios, particularly those involving pedestrian accidents. Using multi-view videos enriched with 3D gaze labels, participants generated structured descriptions and answered reasoning questions. Evaluation combined traditional natural language processing (NLP) metrics with a Large Language Model (LLM)-based semantic scorer.
Warehouse Spatial Intelligence (Track 3): Introducing a novel benchmark, this track focused on fine-grained spatial reasoning in dynamic warehouse environments. AI systems were required to interpret RGB-D inputs and answer spatial questions that combined perception, geometry, and language. Tasks included measuring object distances, assessing global layouts, and responding to natural language spatial queries, with datasets also generated in NVIDIA Omniverse.
Road Object Detection in Fish-Eye Cameras (Track 4): This track emphasized efficient road object detection from fisheye cameras, crucial for traffic monitoring due to their panoramic coverage but challenging due to image distortion. The goal was to support lightweight, real-time deployment on edge devices like NVIDIA Jetson. Submissions were evaluated based on a harmonic mean of F1-score and normalized frame rate, requiring at least 10 frames per second (FPS) on the Jetson AGX Orin 64GB device.
Datasets and Evaluation
The challenge leveraged several specialized datasets. Track 1 utilized the Physical AI Smart Spaces Dataset, a large-scale synthetic dataset with over 250 hours of video from nearly 1,500 indoor cameras. Track 2 employed the Woven Traffic Safety (WTS) dataset, featuring over 1,200 staged and 4,800 real-world pedestrian-related traffic videos, enhanced with 2D and 3D gaze annotations. Track 3 used the Physical AI Spatial Intelligence Warehouse Dataset, containing around 500,000 VQA samples generated through Omniverse simulations. For Track 4, the FishEye8K dataset was used for training and validation, with an in-house test set for final evaluation.
The evaluation framework was designed for fairness and reproducibility, enforcing submission limits and using partially held-out test sets. Teams vying for top rankings were required to publicly release their code, ensuring transparency and contributing to the broader research community.
Also Read:
- Mapping the Road Ahead: Cross-View Transformers for Autonomous Vehicle Perception
- A New Framework for Evaluating AI Training Data Trustworthiness
Leading Solutions and Future Directions
Across all tracks, top-performing teams showcased innovative approaches. In Track 1, winning strategies included offline geometry-centric pipelines fusing depth maps and online tracking systems. Track 2 saw success with large Vision-Language Models (VLMs) enhanced by spatiotemporal prompt engineering. For Track 3, leading solutions combined LLMs with tool-augmented spatial APIs and lightweight architectures. In Track 4, teams excelled with data-centric strategies, lightweight models, and distortion-aware augmentations optimized for edge deployment.
The 9th AI City Challenge highlighted a strong trend towards multimodal fusion, domain-specific model design, and real-time readiness. As AI systems become more integrated into safety-critical scenarios, the continued convergence of perception, language, and reasoning will be crucial for driving innovation in intelligent cities. For more in-depth information, you can refer to the research paper.


