spot_img
HomeResearch & DevelopmentEvaluating Audio Tagging Models on Raspberry Pi: Performance and...

Evaluating Audio Tagging Models on Raspberry Pi: Performance and Thermal Insights

TLDR: This research comprehensively evaluates various CNN-based audio tagging models, including PANNs, ConvNeXt, and MobileNetV3 architectures, on a Raspberry Pi. The study assesses inference time and CPU temperature over 24-hour continuous sessions, comparing performance in headless and GUI-enabled scenarios. Findings reveal that while lightweight models can maintain stable performance, the presence of a graphical user interface significantly increases inference times and CPU temperatures, often to critical levels. The paper provides crucial insights for deploying efficient and thermally stable audio tagging solutions on resource-constrained edge devices.

The deployment of advanced audio tagging models on small, low-power devices like the Raspberry Pi is a growing area of research, crucial for a wide range of real-time applications. From monitoring the daily routines of elderly individuals to tracking wildlife and detecting mechanical failures in industrial settings, these applications often demand immediate processing at the source, a concept known as edge computing. This approach is vital for reducing latency, preserving privacy, and conserving energy, especially where cloud-based processing is impractical.

Previous studies have explored the feasibility of using Convolutional Neural Networks (CNNs) for audio tagging on devices such as the Raspberry Pi. However, these efforts often focused on a single model and encountered challenges related to thermal management and the stability of inference latency over extended periods. This new research takes a more comprehensive approach, evaluating a diverse array of CNN architectures.

The study includes all 1D and 2D models from the Pretrained Audio Neural Networks (PANNs) framework, a ConvNeXt-based model adapted for audio classification, and various MobileNetV3 architectures. Additionally, two recently proposed PANNs-derived networks, CNN9 and CNN13, were also assessed. To ensure efficient deployment and portability across different hardware platforms, all models were converted to the Open Neural Network Exchange (ONNX) format.

A significant differentiator of this research is its extensive evaluation period: each model was subjected to continuous inference tasks for 24 hours. This rigorous testing allowed researchers to meticulously monitor and analyze critical performance metrics, including the stability of inference time, fluctuations in CPU temperature, and overall system reliability. The experiments were conducted using a Raspberry Pi 4B with 4 GB of RAM, an external UGREEN USB 2.0 sound card, and a RODE Lavalier II microphone. The setup also featured a 7-inch LCD touchscreen for the graphical user interface (GUI) and a PiJuice HAT UPS for power stability.

The study compared model performance under two distinct scenarios: a ‘headless’ mode, where the application ran without a visual interface, and a mode where a real-time GUI was active. This comparison was essential for understanding the practical impact of user interfaces on model performance in real-world deployment.

Key Findings on Performance and Thermal Behavior

The results indicate that with appropriate model selection and optimization, it is indeed possible to maintain consistent inference latency and manage thermal behavior effectively over prolonged periods. In the headless scenario, most models demonstrated stable and relatively low inference times. Lightweight models like CNN9, CNN13, and CNN6 consistently showed inference times under or around 1 second. More complex architectures, such as ConvNeXt and ResNet54, exhibited higher but still stable inference times, generally between 2 and 3 seconds.

However, when the GUI was active, a general increase in inference time was observed across all models, accompanied by higher temporal variability. Notably, ConvNeXt, ResNet54, and Res1dNet51 showed peaks exceeding 4 and 5 seconds, which could compromise their suitability for real-time applications. This suggests that the GUI introduces a significant system load that affects models unevenly, with deeper or more computationally intensive architectures experiencing more pronounced degradation.

Regarding CPU temperature, lightweight architectures like MobileNetV2, CNN6, and MobileNetV1 maintained the lowest and most stable temperatures, well below 65°C, in the headless setup. Conversely, heavier models such as ResNet54, ConvNeXt, and Wavegram reached temperatures above 80°C. When the GUI was active, nearly all models converged to higher and more uniform temperature values, clustering around 83–85°C. This effect was particularly evident for models that previously had moderate temperatures in the headless setup, showing a significant increase. Sustained high temperatures can lead to thermal throttling or hardware degradation over long-term operation.

Interestingly, the MobileNet models, which were among the most thermally efficient in headless mode, became some of the highest temperature generators when the GUI was enabled. This suggests that the original PANNs implementation might not be well-suited for scenarios involving a real-time graphical interface.

Also Read:

Implications for Edge AI Solutions

This research offers valuable insights for deploying audio tagging models in real-world edge computing environments. It underscores the importance of considering the full deployment context, including interface requirements, when selecting and validating neural network models for edge applications. Model selection should not only optimize for inference accuracy and latency but also account for thermal constraints and the potential overhead introduced by concurrent system tasks.

For robust and safe deployments, especially in unattended or industrial environments, thermal profiling and system monitoring are essential components of the design and evaluation process. The study concludes that MobileNet-based models, particularly lightweight MobileNetV3 variants, provide a favorable balance between performance and computational efficiency for continuous real-time audio inference on low-power embedded devices. However, the presence of a real-time graphical user interface introduces significant thermal and computational overhead.

Future work aims to restructure the inference pipeline of PANNs models to enhance their performance and thermal efficiency in GUI-active scenarios. Additionally, experiments will be expanded to incorporate other IoT devices, such as NVIDIA Jetson platforms, or specialized AI hardware to further improve performance and scalability.

This comprehensive analysis represents a significant advancement in understanding the practical considerations for deploying various CNN-based audio tagging models on edge devices, facilitating informed decisions for real-world applications in areas like environmental monitoring, smart home systems, and assistive technologies. You can read the full research paper here: Comprehensive Evaluation of CNN-Based Audio Tagging Models on Resource-Constrained Devices.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -