SONG L F, HU Y Y, SONG M F, et al. Action–masked reinforcement learning with decision-tree distillation for ship anti-missile self-defenseJ. Chinese Journal of Ship Research, 2026, 22(X): 1–11 (in Chinese). DOI: 10.19693/j.issn.1673-3185.05130
Citation: SONG L F, HU Y Y, SONG M F, et al. Action–masked reinforcement learning with decision-tree distillation for ship anti-missile self-defenseJ. Chinese Journal of Ship Research, 2026, 22(X): 1–11 (in Chinese). DOI: 10.19693/j.issn.1673-3185.05130

Action–masked reinforcement learning with decision-tree distillation for ship anti-missile self-defense

  • Objective Shipborne anti-missile self-defense involves rapid sequential decision-making under multiple incoming threats, short engagement windows, heterogeneous weapon ranges, limited ammunition, and strict execution constraints. Conventional deep reinforcement learning can learn adaptive interception policies, but unrestricted exploration may generate many infeasible actions, while neural-network policies are difficult to inspect and audit. This study proposes an action-masked proximal policy optimization method with decision-tree policy distillation to improve action executability, training efficiency, and policy interpretability.
    Method A red-blue confrontation scenario with one high-value manned vessel and four unmanned vessels is modeled as a Markov decision process. The defense action is decomposed into platform selection, target selection, weapon selection, and launch quantity. Platform availability, firing cooldown, target validity, repeated-interception status, weapon range, ammunition availability, and launch-quantity limits are encoded as state-dependent masks derived from an expert rule base. The masks are applied hierarchically during both policy sampling and policy updating so that PPO optimizes only over the currently legal action set; when no feasible action exists at a decision layer, the system falls back to a no-launch action. After training, the AM-PPO policy is used as a teacher to generate state-action samples. A CART decision tree is then trained with threat-, resource-, and constraint-related interpretable features to extract explicit decision rules.
    Results Five independent training runs show that the average hit rate over the final 100 episodes increases from 0.3844 for standard PPO to 0.7649 for AM-PPO, corresponding to an improvement of 38.05 percentage points. The between-run standard deviation decreases from 0.1237 to 0.0449, indicating substantially improved training stability. Under the sliding-window convergence criterion, AM-PPO enters the stable convergence stage at episode 148, whereas PPO converges at episode 237. Detailed comparison of a representative run further shows a post-convergence average hit rate of 0.8138 for AM-PPO versus 0.5820 for PPO, while ineffective interception launches are markedly reduced because infeasible platform-target-weapon combinations are removed before sampling. The distilled decision-tree policy achieves a hit rate of 0.8551 in the same evaluation scenario, retaining 88.06% of the performance of the AM-PPO teacher policy with a hit rate of 0.9710.
    Conclusion Embedding explicit engagement constraints into PPO can reduce invalid exploration and improve both convergence efficiency and interception performance, while decision-tree distillation converts the learned neural policy into auditable rules with limited performance degradation. The proposed framework provides a practical route toward executable, reviewable, and interpretable autonomous interception decision-making for highly dynamic and strongly constrained shipborne anti-missile defense missions.
  • loading

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return