
적외선 영상에서의 표적 탐지 성능 향상을 위한 개선된 YOLOX 모델
Ⓒ 2026 Korea Society for Naval Science & Technology
초록
YOLOX는 일반 객체 검출에서 강력한 성능을 보여주지만, neck 구조가 주로 P3–P5 수준의 저해상도 특성 맵에 의존하고 있어 작은 대상에 대한 검출 능력이 제한된다. 본 논문에서는 단순한 구조와 높은 추론 효율성을 갖춘 YOLOX를 베이스라인 모델로 선택했으며, neck 구조에 저수준 C2 특성 및 고수준 의미적 특징을 결합하는 방법을 제안한다. C2 특징이 제공하는 풍부한 공간 정보를 활용함으로써 거짓 양성을 효과적으로 감소시켰다. 제안한 수정된 구조는 기존 YOLOX 대비 mAP50:95에서 최대 7.4%가 향상되었다.
Abstract
Although YOLOX has demonstrated strong performance on generic object detection, its neck relies primarily on the low‑resolution feature maps from levels P3–P5, which limits its ability to detect small targets. In this study, YOLOX was selected as the baseline model, and we propose a method which incorporates the low-level C2 features and high-level semantic features within the neck structure. By exploiting the richer spatial details of the C2 features, we effectively reduce false positives. The proposed modification improves detection performance by up to 7.4 % in mAP50:95 compared to the baseline YOLOX.
Keywords:
Object Detection, Infrared Image, Feature Pyramid, Deep Learning, Convolutional Neural Network키워드:
객체 탐지, 적외선 영상, 특징 피라미드, 딥러닝, 합성곱 신경망References
-
Joseph Redmon, Santosh Divvala, Ross Girshick, & Ali Farhadi, ‘You Only Look Once: Unified, Real-Time Object Detection,’ in proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 779-788.
[https://doi.org/10.1109/CVPR.2016.91]
-
Joseph Redmon & Ali Farhadi, ‘YOLO9000: Better, Faster, Stronger,’ in proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 7263-7271.
[https://doi.org/10.1109/CVPR.2017.690]
- Joseph Redmon & Ali Farhadi, ‘YOLOv3: An Incremental Improvement,’ arXiv preprint arXiv:1804.02767, , 2018.
- Alexey Bochkovskiy, Chien-Yao Wang, & Hong-Yuan Mark Liao, ‘YOLOv4: Optimal Speed and Accuracy of Object Detection,’ arXiv preprint arXiv:2004.10934, , 2020.
- Glenn Jocher, Ayush Chaurasia, Alex Stoken, Jirka Borovec, Yonghye Kwon, Kalen Michael, Jiacong Fang, Lorna Imyhxy, Colin Wong, Yifu Zeng, et al. (2022). ‘ultralytics/yolov5: v6. 2-YOLOv5 Classification Models, Apple M1, Reproducibility, ClearML and Deci.ai Integrations,’ Zenodo.
-
Chien-Yao Wang, Alexey Bochkovskiy, & Hong-Yuan Mark Liao, ‘YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors,’ in proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 7464-7475.
[https://doi.org/10.1109/CVPR52729.2023.00721]
-
Rejin Varghese & M. Sambath, ‘YOLOv8: A Novel Object Detection Algorithm with Enhanced Performance and Robustness,’ in proceedings of 2024 International Conference on Advances in Data Engineering and Intelligent Computing Systems (ADICS), Chennai, India, April 2024, pp. 1-6.
[https://doi.org/10.1109/ADICS58448.2024.10533619]
-
Chien-Yao Wang, I-Hau Yeh, & Hong-Yuan Mark Liao, YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information, In Computer Vision - ECCV 2024, Springer Nature Switzerland, September 2024, pp. 1-21.
[https://doi.org/10.1007/978-3-031-72751-1_1]
- Zheng Ge, Songtao Liu, Feng Wang, Zeming Li, & Jian Sun, ‘YOLOX: Exceeding YOLO Series in 2021,’ arXiv preprint arXiv:2107.08430, , 2021.
- Hao Zhang, Feng Li, Shilong Liu, Lei Zhang, Hang Su, Jun Zhu, Lionel M. Ni, & Heung-Yeung Shum, ‘DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection,’ arXiv preprint arXiv:2203.03605, , 2022.
-
Yian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei, Guanzhong Wang, Qingqing Dang, Yi Liu, & Jie Chen, ‘DETRs Beat YOLOs on Real-Time Object Detection,’ in proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 16965-16974.
[https://doi.org/10.1109/CVPR52733.2024.01605]
- Mingxin Zhao, Li Cheng, Xu Yang, Peng Feng, Liyuan Liu, & Nanjian Wu, ‘TBC-Net: A Real-Time Detector for Infrared Small Target Detection Using Semantic Constraint,’ arXiv preprint arXiv:2001.05852, , 2019.
- Zhiheng Hu, Yongzhen Wang, Peng Li, Jie Qin, Haoran Xie, & Mingqiang Wei, ‘ISmallNet: Densely Nested Network with Label Decoupling for Infrared Small Target Detection,’ in proceedings of ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, June 2023.
-
Xingang Mou, Shuai Lei, & Xiao Zhou, ‘YOLO-FR: A YOLOv5 Infrared Small Target Detection Algorithm Based on Feature Reassembly Sampling Method,’ Sensors, VOL. 23, NO. 5, 2023, Article 2710.
[https://doi.org/10.3390/s23052710]
-
Huicong Wang, Kaijun Ma, Juan Yue, Yuhan Li, Jiaxin Huang, Jie Liu, Linhan Li, Xiaoyu Wang, Nengbin Cai, & Sili Gao, ‘Small-Target Detection Based on Improved YOLOv8 for Infrared Imagery,’ Electronics, VOL. 14, NO. 5, 2025, Article 947.
[https://doi.org/10.3390/electronics14050947]
-
Chao Wang, Rongdi Wang, Ziwei Wu, Zetao Bian, & Tao Huang, ‘YOLO-UIR: A Lightweight and Accurate Infrared Object Detection Network Using UAV Platforms,’ Drones, VOL. 9, NO. 7, 2025, 479.
[https://doi.org/10.3390/drones9070479]
-
Nan Jiang, Kuiran Wang, Xiaoke Peng, Xuehui Yu, Qiang Wang, Junliang Xing, Guorong Li, Jian Zhao, Guodong Guo, & Zhenjun Han, ‘Anti-UAV: A Large-Scale Benchmark for Vision-Based UAV Tracking,’ IEEE Transactions on Multimedia, VOL. 25, November 2021, pp. 486-500.
[https://doi.org/10.1109/TMM.2021.3128047]
-
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, & Dhruv Batra, ‘Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization,’ in proceedings of the IEEE International Conference on Computer Vision, 2017, pp. 618-626.
[https://doi.org/10.1109/ICCV.2017.74]