Cooperative path planning for heterogeneous USV swarms based on multi-agent reinforcement learning and hierarchical optimization
-
Abstract
Objectives To investigate cooperative path planning for heterogeneous unmanned surface vehicle (USV) swarms under simulation settings with analytical currents and simplified moving threats, while accounting for platform-level kinematic differences and coupled constraints, a cooperative path planning method integrating multi-agent reinforcement learning with hierarchical optimization (MARL-BiOpt) is proposed. Methods A hierarchical optimization framework is constructed: the upper level employs an attention-enhanced multi-agent proximal policy optimization (Att-MAPPO) algorithm that generates cooperative heading and velocity strategies for heterogeneous platforms through global situational awareness; the lower level embeds a sequential least squares programming (SLSQP) solver to locally refine the waypoints under hard constraints including heterogeneous kinematics, safety distances, and ocean current compensation. A staged curriculum learning strategy is designed to accelerate convergence, and a refinement-offset-based reward feedback mechanism is proposed to achieve coordinated optimization between the two levels. Results In simulation experiments across three typical maritime task scenarios, the MARL-BiOpt method reduces the total path length by an average of 13.1% and the three-scenario mean constraint violation rate to approximately 0.23% compared to pure MAPPO; compared with the RRT*+SLSQP method, the planning time is reduced by an average of 72.8% and the dynamic adaptation success rate is improved by 23.6 percentage points. Conclusions The proposed method effectively balances global optimality and constraint feasibility through the integration of data-driven and model-driven paradigms, providing a methodological reference for heterogeneous unmanned-system cooperation under the specified simulation assumptions.
-
-