spot_img
HomeResearch & DevelopmentYOLOatr: Advancing Automatic Target Recognition in Challenging Thermal Infrared...

YOLOatr: Advancing Automatic Target Recognition in Challenging Thermal Infrared Environments

TLDR: YOLOatr is a novel deep learning model based on a modified YOLOv5s, specifically designed for Automatic Target Detection (ATD) and Recognition (ATR) in challenging thermal infrared (TI) imagery for defense and surveillance. It addresses common issues like limited datasets, scale variations, and occlusion through architectural enhancements (extra small detection head, BiFPN neck) and a custom data augmentation profile. Evaluated on the DSIAC MWIR dataset, YOLOatr achieved up to 99.6% accuracy on correlated test ranges and demonstrated a significant 11.4% performance gain over baseline YOLOv5s on decorrelated ranges, proving its superior generalization ability and state-of-the-art performance for real-time, high-accuracy target recognition in complex thermal environments.

Automatic Target Detection (ATD) and Recognition (ATR) in thermal infrared (TI) imagery is a crucial yet challenging task, particularly in defense and surveillance applications. Unlike commercial autonomous vehicle systems, this domain faces unique hurdles such as limited datasets, issues with scale invariance over long distances, deliberate target occlusion, low sensor resolution leading to a lack of structural information, and the impact of varying weather, temperature, and time of day. These factors contribute to increased variability within target classes and greater similarity between different classes, making accurate real-time ATR difficult.

Introducing YOLOatr: A Specialized Approach

To address these challenges, researchers have proposed YOLOatr, a modified anchor-based single-stage detector. This innovative model is built upon a modified version of YOLOv5s, a popular deep learning architecture known for its balance of speed and accuracy. The key enhancements in YOLOatr include optimized modifications to its detection heads, improved feature-fusion in the network’s neck, and a custom data augmentation profile tailored for thermal imagery.

The development of YOLOatr focused on creating a robust solution capable of reliably detecting small infrared targets in highly cluttered environments, across multiple distances, and under various illumination and signal-to-clutter noise ratio (SCNR) conditions (both day and night).

Methodology and Dataset

The development process for YOLOatr involved two main phases: extensive preprocessing of the dataset and an experimental approach with an ablation study on various YOLOv5s model variants. The goal was to identify the optimal modifications in terms of training methodology, architecture, and data augmentation.

The model was evaluated using the comprehensive DSIAC MWIR dataset, which is the largest publicly available thermal infrared dataset for ground-based tactical platforms. This dataset includes images of 13 different tactical and civilian vehicles captured at distances ranging from 1000m to 5000m, with data collected during both day and night. For their experiments, the researchers focused on four specific vehicles: T72, BRDM2, Pickup, and SUV. The dataset was partitioned using two protocols: ‘correlated dataset partitioning’ (T1), where training and testing images came from the same range, and ‘decorrelated dataset partitioning’ (T2), where training occurred at lower ranges and testing at higher ranges to assess generalization ability.

Architectural Innovations and Learning

Given that thermal infrared ATR demands high accuracy and real-time inference speed, YOLOv5s was chosen as the base model. A significant architectural modification was the introduction of an extra small detection head (P2) to improve the detection of small objects, a known challenge for YOLOv5. Additionally, the original PANet neck was replaced with a BiFPN (Weighted Bi-directional Feature Pyramid) Network neck, designed to enhance feature fusion and propagation within the network.

The researchers also experimented with different learning methodologies, finding that training the model from scratch yielded slightly better results than transfer learning using ImageNet pre-trained weights. This suggests that features learned from RGB images may not align well with the unique characteristics of thermal infrared objects.

A custom augmentation profile (CAP) was developed to diversify input images during training and improve model accuracy. This profile carefully adjusted parameters for brightness, contrast, translation, rotation, scaling, flipping, and perspective changes, while reducing or disabling augmentations like shear and mosaic that could be counterproductive for already low-resolution thermal images.

Performance and Results

YOLOatr demonstrated impressive performance. When tested on the correlated dataset (T1), it achieved a mean Average Precision (mAP) of 99.6%, showing a slight but notable gain over the baseline YOLOv5s model (99.4%). This highlights the effectiveness of the single-frame, single-stage, anchor-based deep convolutional neural network for high-accuracy detection with a lean structure.

More significantly, YOLOatr showcased superior generalization ability on the decorrelated dataset (T2). While the baseline YOLOv5s model’s accuracy declined significantly to 27.1% mAP on T2, YOLOatr achieved a mAP of 37.7%, representing an 11.4% performance gain. This improvement is crucial as it indicates the model’s ability to perform well even when tested on ranges different from its training data, where targets appear smaller and clutter changes. The T-72 Tank, a tactical vehicle, showed the highest accuracy at T2 (62.2%), likely due to its size and thermal signature compared to commercial vehicles like the Pickup truck (16.3% mAP).

Overall, YOLOatr achieves state-of-the-art ATR performance, with near-perfect accuracy for small objects over long ranges in highly cluttered infrared images. It is a lean architecture with a very fast inference speed of 110 frames per second (fps), making it a viable detector for real-time applications with limited computational requirements.

Also Read:

Future Directions

While YOLOatr represents a significant advancement, the researchers acknowledge certain limitations and areas for future work. These include exploring optimum hyperparameter selection via genetic algorithms, conducting a comparative analysis of model accuracy with False Alarm Rate (FAR), and investigating further modifications to the head structure for improved feature extraction. Future research will also involve evaluating YOLOatr’s performance against more recently proposed YOLO variants like YOLOv7, YOLOv8, and YOLO-NAS.

For more in-depth information, you can access the full research paper here.

Nikhil Patel
Nikhil Patelhttps://blogs.edgentiq.com
Nikhil Patel is a tech analyst and AI news reporter who brings a practitioner's perspective to every article. With prior experience working at an AI startup, he decodes the business mechanics behind product innovations, funding trends, and partnerships in the GenAI space. Nikhil's insights are sharp, forward-looking, and trusted by insiders and newcomers alike. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -