New AI Model Sets Record for Detecting Tiny Infrared Targets, Boosting Wildfire and Surveillance Capabilities

By The Building Texas Show•
Researchers at Harbin Institute of Technology have developed a diffusion-enhanced Mamba network that achieves state-of-the-art infrared small target detection, with potential applications in wildfire prevention, surveillance, and autonomous navigation.

Found this article helpful?

Share it with your network and spread the knowledge!

New AI Model Sets Record for Detecting Tiny Infrared Targets, Boosting Wildfire and Surveillance Capabilities

Infrared small target detection is essential for remote sensing, fire prevention, and surveillance, but existing methods often struggle with tiny, low-contrast targets that lack distinct shape or texture. A new two-stage deep learning network developed by researchers at the Research Center for Space Optical Engineering at Harbin Institute of Technology combines diffusion-based feature enhancement with state-space modeling to suppress background clutter and amplify target signals. The approach, published on June 30, 2026, in the Journal of Remote Sensing, achieves state-of-the-art detection rates across three public datasets, significantly reducing both missed detections and false alarms in complex imaging environments.

The technology directly impacts forest fire prevention, surveillance early warning systems, and military threat assessment—applications where missed detections or false alarms can have severe consequences. Infrared targets often occupy fewer than 81 pixels (typically under 9×9), exhibit extremely low energy with signal-to-noise ratios around 3, and lack prominent shape or texture information, causing them to be easily submerged in background clutter. Deep learning methods have improved performance, but most focus exclusively on target features while neglecting background information, leading to severe class imbalance. The proposed Diffusion-Enhanced Dense Mamba Network (DEDM-Net) addresses this by simultaneously modeling both targets and backgrounds.

The two-stage network achieves a synergistic effect. The first stage employs a dual-path diffusion model with a novel blind processing module that predicts each pixel using only surrounding information—never the pixel itself—preventing extremely small targets from being misclassified as background. The second stage introduces a dense nested Mamba architecture based on the state space model, which captures long-range correlations across global and local features with linear computational complexity—a significant advantage over conventional Transformers. A cross-stage prediction fusion module further integrates features from both stages, improving contour segmentation accuracy. Together, these innovations deliver superior performance across all evaluation metrics compared to 11 state-of-the-art methods.

DEDM-Net was evaluated on three public datasets: NUAA-SIRST (427 images), NUDT-SIRST (1,327 images at 256×256), and IRSTD-1k (1,000 images at 512×512). On the NUDT-SIRST dataset, the method achieved 93.40% IoU, 93.28% nIoU, 98.37% detection probability, and a remarkably low false-alarm rate of just 3.75×10⁻⁶—outperforming DNA-Net (92.99% IoU, 93.22% nIoU) and ISTDU-Net (91.69% IoU, 91.84% nIoU). On the IRSTD-1k dataset, DEDM-Net achieved 73.71% IoU and 93.89% detection probability with only 11.10×10⁻⁶ false alarms, surpassing all competitors.

"Infrared small targets are extremely challenging because they lack shape and texture—they're essentially just a few bright pixels in a sea of background," said corresponding author Dr. Shikai Jiang of Harbin Institute of Technology. "By modeling both the target-free background and potential target regions simultaneously, our diffusion-enhanced approach effectively amplifies what matters while suppressing what doesn't. The Mamba architecture then provides the global context needed to distinguish true targets from bright clutter."

While DEDM-Net achieves superior accuracy, the diffusion-based two-stage design increases inference time compared to single-stage networks. Future work will focus on model distillation, mixed-precision inference, and faster samplers to reduce the required diffusion steps. The approach holds promise for real-time surveillance systems, autonomous drone navigation in low-visibility conditions, and early wildfire detection networks. The framework could also inspire new thinking about how generative models and state-space architectures can be combined for other challenging computer vision tasks where target-background separation is critical. The full study is available at DOI: 10.34133/remotesensing.1046.