Action-masked reinforcement learning with decision-tree distillation for ship air defense
-
Abstract
Objectives To address the problems of numerous incoming targets, short engagement windows, limited weapon resources, and the poor interpretability of deep reinforcement learning policies in shipborne air self-defense missions, this study investigates an autonomous interception decision-making method that balances action executability and policy credibility.Methods Based on a red-blue confrontation simulation environment, the shipborne air self-defense process is modeled as a Markov decision process, and a Proximal Policy Optimization (PPO) algorithm based on deep reinforcement learning is adopted to learn air-defense engagement policies; furthermore, procedural knowledge and decision-making knowledge are transformed into state-dependent action masks to suppress invalid and illegal actions, and the trained neural-network policy is distilled by means of a decision tree. Results Prior knowledge is transformed into a constrained action space to enable efficient PPO-based policy learning, and policy interpretability is achieved through decision-tree rule extraction. ConclusionsThe proposed method provides an implementation pathway for executable and auditable autonomous decision-making in highly dynamic and strongly constrained anti-missile scenarios.
-
-