
군집 USV 무장할당을 위한 다중 에이전트 강화학습 기반 실시간 분산 의사결정 구조 연구
Ⓒ 2026 Korea Society for Naval Science & Technology
초록
본 논문은 전투용 군집 무인수상정의 교전 효율 극대화를 위한 다중 에이전트 강화 학습 기반 실시간 분산 의사결정 구조를 제안한다. 이는 무기표적할당 문제를 시간윈도우가 포함된 다중모드 자원제약 프로젝트 스케줄링 문제로 변환한 후 실시간 최적화를 위해 리스트스케줄링 휴리스틱으로 초기해를 찾고, 가변이웃탐색으로 해를 개선한다. 특히, 그래프 신경망으로 동적 전장 상황을 임베딩하고, 중앙집중식 학습 및 분산형 실행 구조의 근접 정책 최적화 알고리즘을 융합하여 데이터링크 소실 같은 통신단절 상황에서도 에이전트의 자율적 전술 결심이 가능하게 한다. 시뮬레이션 결과 5초의 임계 시간 내에 근사 최적해를 안정적으로 산출하며, 기존 방식 대비 생존성과 표적 무력화율을 향상시켰다.
Abstract
This study proposes a multi-agent reinforcement learning framework to maximize the engagement efficiency of combat unmanned surface vehicle swarms. The weapon target assignment problem is formulated as a multi-mode resource-constrained project scheduling problem with time windows. For real-time optimization, a list-scheduling heuristic is used to generate initial solutions, which are refined by variable neighborhood search. In particular, graph neural networks are used to represent dynamic topologies, and a centralized training and decentralized execution (CTDE)-based multi-agent proximal policy optimization framework enables autonomous decision-making during data link loss. Simulation results demonstrate that the proposed model achieves near-optimal solutions within 5 s, thereby improving survivability and target neutralization.
Keywords:
Weapon Target Assignment, Multi-Agent Reinforcement Learning, Variable Neighborhood Search, Graph Neural Networks, Multi-Agent Proximal Policy Optimization키워드:
무기표적할당, 다중 에이전트 강화학습, 가변이웃탐색, 그래프 신경망, 다중 에이전트 근접 정책 최적화Acknowledgments
이 논문은 2026년 정부(방위사업청)의 재원으로 국방기술진흥연구소의 지원을 받아 수행된 연구임(KRIT-CT-25-011*)
References
- DARPA, Mosaic Warfare, DARPA News, 2018. https://www.darpa.mil/news/mosaic-warfare, (accessed 2026.04.01.)
-
Jinsung Lee & Jongchul Na, ‘Required Capabilities for Maritime Manned-Unmanned Teaming,’ Journal of KNST, VOL. 6, NO. 3, 2023, pp. 308-313.
[https://doi.org/10.31818/JKNST.2023.09.6.3.308]
-
Sang-Kyum Na & Dong-Sun Park, ‘A Study on the Autonomous Level of the Maritime Manned and Unmanned Combined Combat System,’ Journal of KNST, VOL. 6, NO. 3, 2023, pp. 286-292.
[https://doi.org/10.31818/JKNST.2023.09.6.3.286]
- 한종환, ‘러-우 전쟁 해양작전 교훈과 한국 해군의 전략/전력 발전에 대한 함의,’ 국방연구, 제68권 제2호, 2025, pp. 157-186.
-
Tao Hu, Xiaoxue Zhang, Xueshan Luo, & Tao Chen, ‘Dynamic Target Assignment by Unmanned Surface Vehicles Based on Reinforcement Learning,’ Mathematics, VOL. 12, NO. 16, 2024, Article 2557.
[https://doi.org/10.3390/math12162557]
-
Chan-Woo Kim, Seong-Hyeon Ju, & Kyung-Min Seo, ‘Discrete-Event Simulation Framework for Composite Warfare with Modular CFCS Modeling for Naval Combat System,’ IEEE Access, VOL. 14, 2025, pp. 37-53.
[https://doi.org/10.1109/ACCESS.2025.3648783]
- 이우석, ‘한국형 통합 군사드론 체계의 발전을 위한 제언,’ 합참지, VOL. 99, 2024, pp. 69-81.
-
Woo Seok Lee, Jae-Kwan Ryu, JooYoung Lee, & Bongwan Choi, ‘A Study on Integrated Command and Control Methods for Combat Unmanned Surface Vehicle Swarm Operations,’ Journal of KNST, VOL. 8, NO. 4, 2025, pp. 816-832.
[https://doi.org/10.31818/JKNST.2025.12.8.4.816]
-
Guoqing Xia, Xianxin Sun, & Xiaoming Xia, ‘Distributed Swarm Control Algorithm of Multiple Unmanned Surface Vehicles Based on Grouping Method,’ Journal of Marine Science and Engineering, VOL. 9, NO. 12, 2021, Article 1324.
[https://doi.org/10.3390/jmse9121324]
-
Manuel Blanco Abello & Zbigniew Michalewicz, ‘Multiobjective Resource-Constrained Project Scheduling with a Time-Varying Number of Tasks,’ The Scientific World Journal, 2014, Article 420101.
[https://doi.org/10.1155/2014/420101]
- Woo Seok Lee, Jae-deok Jang, Bong Wan Choi, & JooYoung Lee, ‘Missile Allocation Using Multi-Mode Resource-Constrained Project Scheduling Problem with Time Window,’ Journal of the Korea Society of Systems Engineering, VOL. 21, 2025, pp. 63-77.
-
Ling Wu, Changfeng Xing, Faxing Lu, & Peifa Ha, ‘An Anytime Algorithm Applied to Dynamic Weapon-Target Allocation Problem with Decreasing Weapons and Targets,’ in proceedings of IEEE Congress on Evolutionary Computation, June 2008, pp. 1703-1710.
[https://doi.org/10.1109/CEC.2008.4631306]
-
Yannick Molinghen, Augustin Delecluse, Renaud De Landtsheer, & Stefano Michelini, ‘Reinforcement Learning Methods for Neighborhood Selection in Local Search,’ arXiv preprint arXiv:2601.07948, , 2026.
[https://doi.org/10.5220/0014445200004055]
-
Sinki Jeong, Hyunseop Uhm, & Young Hoon Lee, ‘Rolling-Horizon Scheduling Algorithm for Dynamic Weapon-Target Assignment in Air Defense Engagement,’ Journal of the Korean Institute of Industrial Engineers, VOL. 46, NO. 1, 2020, pp. 11-24.
[https://doi.org/10.7232/JKIIE.2020.46.1.011]
-
Anthony Goeckner, Yueyuan Sui, Nicolas Martinet, Xinliang Li, & Qi Zhu, ‘Graph Neural Network-Based Multi-Agent Reinforcement Learning for Resilient Distributed Coordination of Multi-Robot Systems,’ arXiv preprint arXiv:2403.13093, , 2024.
[https://doi.org/10.1109/IROS58592.2024.10802510]
- Haoran Su, ‘Hierarchical GNN-Based Multi-Agent Learning for Dynamic Queue-Jump Lane and Emergency Vehicle Corridor Formation,’ arXiv preprint arXiv:2601.04177, , 2026.
-
Alan S. Manne, ‘A Target-Assignment Problem,’ Operations Research, VOL. 6, NO. 3, 1958, pp. 346–351.
[https://doi.org/10.1287/opre.6.3.346]
-
Alexandre Colaers Andersen, Konstantin Pavlikov, & Túlio A. M. Toffolo, ‘Weapon-Target Assignment Problem: Exact and Approximate Solution Algorithms,’ Annals of Operations Research, VOL. 312, 2022, pp. 581-606.
[https://doi.org/10.1007/s10479-022-04525-6]
-
Xiangping Zeng, Yunlong Zhu, Lin Nan, Kunyuan Hu, Ben Niu, & Xiaoxian He, ‘Solving Weapon-Target Assignment Problem Using Discrete Particle Swarm Optimization,’ in proceedings of 2006 6th World Congress on Intelligent Control and Automation, June 2006.
[https://doi.org/10.1109/WCICA.2006.1713032]
-
Lingren Kong, Jianzhong Wang, & Peng Zhao, ‘Solving the Dynamic Weapon Target Assignment Problem by an Improved Multiobjective Particle Swarm Optimization Algorithm,’ Applied Sciences, VOL. 11, NO. 19, 2021, Article 9254.
[https://doi.org/10.3390/app11199254]
-
Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, & Yi Wu, ‘The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games,’ Advances in Neural Information Processing Systems, VOL. 35, 2022, pp. 24611-24624.
[https://doi.org/10.52202/068431-1787]
- Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, & Pietro Liò, ‘Graph Attention Networks,’ in proceedings of 6th International Conference on Learning Representations (ICLR), 2018.
-
Sihan Wang & Chengjun Ji, ‘A Reinforcement Learning-Variable Neighborhood Search Method for the Cloud Manufacturing Scheduling Robust Optimization Problem with Uncertain Service Time,’ in proceedings of the 2023 4th International Conference on Management Science and Engineering Management, 2024.
[https://doi.org/10.2991/978-94-6463-256-9_54]
-
Nader Shamami, Esmaeil Mehdizadeh, Mehdi Yazdani, & Farhad Etebari, ‘A Hybrid BA-VNS Algorithm for Solving the Weapon Target Assignment Considering Mobility of Resources,’ Economic Computation and Economic Cybernetics Studies and Research, VOL. 57, 2023, pp. 59-76.
[https://doi.org/10.24818/18423264/57.3.23.04]
-
Binhui Chen, Rong Qu, Ruibin Bai, & Wasakorn Laesanklang, ‘A Variable Neighborhood Search Algorithm with Reinforcement Learning for a Real-Life Periodic Vehicle Routing Problem with Time Windows and Open Routes,’ Rairo Operations Research, VOL. 54, NO. 5, 2020, pp. 1467-1494.
[https://doi.org/10.1051/ro/2019080]
- Alexander G. Kline, Real-Time Heuristic Algorithms for the Static Weapon-Target Assignment Problem, Defense Technical Information Center (DTIC), Tech. Rep. AD1055142, 2018.
- 한국방위산업진흥회, ‘LIG넥스원-국방기술진흥연구소, 전투용 무인수상정 핵심 패키지 기술 협약,’ 국방과 기술, 제563호, 2026. https://www.dbpia.co.kr/journal/articleDetail?nodeId=NODE12554202, (accessed 2026.04.01.)
- LIG넥스원, ‘LIG넥스원, 전투용 무인수상정 핵심 패키지 기술 협약 체결,’ LIG Nex1 회사소식, 2025. https://www.lignex1.com/news/nex1newsView.do?bbs_no=7293, (accessed 2026.04.01.)
- 이영근, ‘[심층 분석] K-해군 '네이비 씨 고스트'의 핵심 퍼즐, LIG넥스원의 유무인 복합전투체계 전략,’ 뉴스밸류, 2026.04.01. https://www.newsvalue.kr/news/articleView.html?idxno=23379, (accessed 2026.04.01.)
- 김은규, ‘LIG넥스원, AI 기반 군집 자폭형 무인기 첫 공개,’ 디일렉, 2026.02.25. https://www.thelec.kr/news/articleView.html?idxno=52683, (accessed 2026.04.01.)
- 김성식, ‘LIG넥스원 "전투용 무인수상정, 핵심은 탑재 무기···당사 직접 제작",’ 뉴스1, 2025.05.31. https://www.news1.kr/industry/general-industry/5800759, (accessed 2026.04.01.)
-
Hui Xie, ‘Cooperative Control of Multi-Agent Systems Under Communication Delays and Packet Loss Scenarios,’ IEEE Access, VOL. 12, 2024, pp. 149804-149813.
[https://doi.org/10.1109/ACCESS.2024.3477713]
-
Guihao Wang, Fengmin Wang, Jiahe Wang, Mengzhen Li, Ling Gai, & Dachuan Xu, ‘Collaborative Target Assignment Problem for Large-Scale UAV Swarm Based on Two-Stage Greedy Auction Algorithm,’ Aerospace Science and Technology, VOL. 149, 2024, Article 109146.
[https://doi.org/10.1016/j.ast.2024.109146]
- 윤영혜, ‘LIG넥스원, 전투용 무인수상정 기술 개발 착수,’ 뉴스토마토, 2025.12.23. https://www.newstomato.com/ReadNews.aspx?no=1285742, (accessed 2026.04.01.)
-
Yang Xiong, Shangwen Wang, Hongjun Tian, Guijie Liu, Zihao Shan, Yijie Yin, Jun Tao, Haonan Ye, & Ying Tang, ‘Spatiotemporal Meta-Reinforcement Learning for Multi-USV Adversarial Games Using aHybrid GAT-Transformer,’ Journal of Marine Science and Engineering, VOL. 13, NO. 8, 2025, Article 1593.
[https://doi.org/10.3390/jmse13081593]