한국해군과학기술학회
[ Article ]
Journal of the KNST - Vol. 9, No. 2, pp.504-522
ISSN: 2635-4926 (Print)
Print publication date 30 Jun 2026
Received 09 Apr 2026 Revised 22 Apr 2026 Accepted 02 Jun 2026
DOI: https://doi.org/10.31818/JKNST.2026.6.9.2.504

군집 USV 무장할당을 위한 다중 에이전트 강화학습 기반 실시간 분산 의사결정 구조 연구

이우석1 ; 김준영2 ; 이혁주1 ; 최봉완3, *
1한남대학교 산업공학과 연구원
2LIG넥스원 무인수상정개발단 수석연구원
3한남대학교 산업공학과 교수
Multi-Agent Reinforcement Learning-Based Real-Time Distributed Decision-Making Framework for Swarm USV Weapon Target Assignment
Woo Seok Lee1 ; Jun Young Kim2 ; Hyuk Joo Lee1 ; Bong Wan Choi3, *
1Researcher, Dept. of Industrial & Management Engineering, Hannam University
2Chief system engineer, Unmanned Surface Vehicle Systems R&D, LIG Nex1
3Professor, Dept. of Industrial & Management Engineering, Hannam University

Correspondence to: *Bong Wan Choi 70 Hannam-ro, Daedeok-gu, Daejeon, 34430, Republic of Korea Tel: +82-42-629-7989, +82-42-629-8530 E-mail: bwchoi721@hnu.kr

Ⓒ 2026 Korea Society for Naval Science & Technology

초록

본 논문은 전투용 군집 무인수상정의 교전 효율 극대화를 위한 다중 에이전트 강화 학습 기반 실시간 분산 의사결정 구조를 제안한다. 이는 무기표적할당 문제를 시간윈도우가 포함된 다중모드 자원제약 프로젝트 스케줄링 문제로 변환한 후 실시간 최적화를 위해 리스트스케줄링 휴리스틱으로 초기해를 찾고, 가변이웃탐색으로 해를 개선한다. 특히, 그래프 신경망으로 동적 전장 상황을 임베딩하고, 중앙집중식 학습 및 분산형 실행 구조의 근접 정책 최적화 알고리즘을 융합하여 데이터링크 소실 같은 통신단절 상황에서도 에이전트의 자율적 전술 결심이 가능하게 한다. 시뮬레이션 결과 5초의 임계 시간 내에 근사 최적해를 안정적으로 산출하며, 기존 방식 대비 생존성과 표적 무력화율을 향상시켰다.

Abstract

This study proposes a multi-agent reinforcement learning framework to maximize the engagement efficiency of combat unmanned surface vehicle swarms. The weapon target assignment problem is formulated as a multi-mode resource-constrained project scheduling problem with time windows. For real-time optimization, a list-scheduling heuristic is used to generate initial solutions, which are refined by variable neighborhood search. In particular, graph neural networks are used to represent dynamic topologies, and a centralized training and decentralized execution (CTDE)-based multi-agent proximal policy optimization framework enables autonomous decision-making during data link loss. Simulation results demonstrate that the proposed model achieves near-optimal solutions within 5 s, thereby improving survivability and target neutralization.

Keywords:

Weapon Target Assignment, Multi-Agent Reinforcement Learning, Variable Neighborhood Search, Graph Neural Networks, Multi-Agent Proximal Policy Optimization

키워드:

무기표적할당, 다중 에이전트 강화학습, 가변이웃탐색, 그래프 신경망, 다중 에이전트 근접 정책 최적화

Acknowledgments

이 논문은 2026년 정부(방위사업청)의 재원으로 국방기술진흥연구소의 지원을 받아 수행된 연구임(KRIT-CT-25-011*)

References