spot_img
HomeResearch & DevelopmentAdvancing Audio-Based Pedestrian Detection Amidst Urban Vehicular Noise

Advancing Audio-Based Pedestrian Detection Amidst Urban Vehicular Noise

TLDR: A new study introduces the ASPED v.b dataset, a 1321-hour roadside audio dataset with vehicular noise, to improve audio-based pedestrian detection. It finds that models trained with vehicular noise generalize better and are less prone to misclassifying vehicle sounds as pedestrians compared to models trained in noise-limited environments. The research highlights the need for domain adaptation and diverse training data for robust urban pedestrian sensing.

A new research paper explores the challenging task of audio-based pedestrian detection, particularly in environments filled with vehicular noise. This study introduces a significant new dataset and provides detailed analysis into how different acoustic environments impact the performance and generalizability of audio-based pedestrian detection models.

Traditionally, pedestrian detection has relied heavily on vision-based systems like cameras. However, these systems face limitations in low-light or visually obstructed areas and often raise privacy concerns. Audio-based sensing offers a promising alternative, being more affordable, energy-efficient, and effective in diverse conditions, while also potentially offering better privacy protection.

The researchers, including Yonghyun Kim, Chaeyeon Han, Akash Sarode, Noah Posner, Subhrajit Guhathakurta, and Alexander Lerch, highlight that previous audio-based pedestrian detection efforts have primarily focused on noise-limited environments. Their work addresses this gap by introducing a comprehensive 1321-hour roadside dataset, named ASPED v.b. This dataset is rich with traffic sounds and includes 16 kHz audio synchronized with frame-level pedestrian annotations and video thumbnails, providing a realistic urban soundscape for model training and evaluation.

The study conducted three main analyses. First, a cross-dataset evaluation assessed how well models trained on one type of environment (noise-limited vs. noisy) performed on the other. Second, they investigated the direct impact of vehicular noise in training data on model performance. Finally, they explored the models’ predictive robustness on sounds outside their primary training domain.

Key findings reveal that models struggle with generalization across different acoustic environments. A model trained in a vehicle-free setting (ASPED v.a) showed a performance drop when tested in an environment with vehicular noise (ASPED v.b), and vice-versa. This indicates that the specific background noise characteristics of the training environment significantly influence a model’s ability to perform in varied real-world scenarios.

Crucially, the research demonstrated that models trained with vehicular noise (using ASPED v.b) were more effective at distinguishing pedestrian sounds from vehicle sounds. In contrast, models trained without traffic noise (using ASPED v.a) frequently misclassified vehicle sounds as pedestrian presence, leading to more false positives. This suggests that exposing models to diverse and realistic noise conditions during training is vital for building robust detection systems.

The study also delved into what acoustic features the models were “listening to.” It found that models do not simply rely on sound energy levels. While human speech-related sounds generally had higher detection probabilities, sounds directly associated with pedestrian movement, such as “walk, footsteps” and “run,” were surprisingly ranked lower among human sound categories. Furthermore, the v.a-trained model showed a higher tendency to misclassify certain non-human sounds, like musical instruments, as pedestrians.

Also Read:

In conclusion, this research underscores the critical role of the acoustic environment in developing effective audio-based pedestrian detection systems. The limited generalization capabilities observed suggest a strong need for future work in domain adaptation techniques. Researchers plan to explore methods to help models filter out irrelevant background noise while maintaining sensitivity to subtle pedestrian cues, and potentially integrate multi-modal information, such as visual cues, to enhance robustness in challenging urban settings. You can read the full research paper for more details here.

Karthik Mehta
Karthik Mehtahttps://blogs.edgentiq.com
Karthik Mehta is a data journalist known for his data-rich, insightful coverage of AI news and developments. Armed with a degree in Data Science from IIT Bombay and years of newsroom experience, Karthik merges storytelling with metrics to surface deeper narratives in AI-related events. His writing cuts through hype, revealing the real-world impact of Generative AI on industries, policy, and society. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -