WANG Y X, ZHANG Y L. Learning-guided cooperative evolution for multi-AUV task allocation in dynamic flow fieldsJ. Chinese Journal of Ship Research, 2026, 21(X): 1–14 (in Chinese). DOI: 10.19693/j.issn.1673-3185.05087
Citation: WANG Y X, ZHANG Y L. Learning-guided cooperative evolution for multi-AUV task allocation in dynamic flow fieldsJ. Chinese Journal of Ship Research, 2026, 21(X): 1–14 (in Chinese). DOI: 10.19693/j.issn.1673-3185.05087

Learning-guided cooperative evolution for multi-AUV task allocation in dynamic flow fields

  • Objective Aiming at the problems of strong environmental uncertainty, dynamic variation of task adaptation relationships, and difficulty in balancing search and convergence in cooperative task allocation for multiple autonomous underwater vehicles (AUVs) under dynamic flow-field environments, a cooperative task allocation method for multi-AUV systems in complex dynamic flow-field conditions is studied.
    Method A dual-feedback-driven learning-guided cooperative evolutionary framework is proposed. The framework integrates three coordinated layers: flow-field uncertainty perception, task-structure guidance, and adaptive evolutionary control. A probabilistic long short-term memory (LSTM) network is employed to predict the probability distribution of navigation time correction factors under flow-field effects, providing the predicted mean and standard deviation to characterize the expected navigation-time effect and its uncertainty, respectively. A risk-aware cost function is then constructed to explicitly incorporate flow-field uncertainty into task allocation evaluation. Based on these risk features, a graph neural network (GNN) integrating flow-field information is designed to learn the structural adaptation relationships between tasks and AUVs. The GNN generates a task-AUV preference matrix and guides the auction-based task assignment by dynamically modifying allocation costs, thereby introducing environmental information directly into task-allocation structure generation. In parallel, the same uncertainty information, together with population evolution states, is fed into a reinforcement learning meta-controller based on proximal policy optimization (PPO). The PPO controller adaptively regulates the global mutation probability and auction-based mutation probability according to environmental risk and evolutionary conditions, enabling a dynamic balance between exploration and exploitation. Consequently, flow-field uncertainty acts on both the structure space and search space, forming a closed-loop optimization process of “LSTM risk perception–GNN structural guidance–PPO search regulation–evolutionary feedback”.
    Results Simulation experiments are conducted in dynamic flow-field environments. The results show that the proposed method reduces the total system energy consumption by 53% and decreases the dynamic flow-field risk penalty by 74% compared with the classical non-dominated sorting genetic algorithm. The proposed method also achieves competitive task completion efficiency while providing higher feasible-solution ratios and improved constraint satisfaction. Mechanism-replacement ablation experiments further demonstrate that uncertainty perception, structural guidance, and adaptive search regulation provide complementary contributions: replacing probabilistic uncertainty perception with mean-only prediction causes the most significant degradation in overall optimization quality, while replacing GNN guidance or PPO control with handcrafted alternatives mainly deteriorates load balancing and energy-control performance, respectively. Scalability experiments further show that the method can maintain effective Pareto-front generation with increasing task and AUV numbers, while the computational cost remains controllable.
    Conclusion The proposed method establishes an explicit pathway for transforming flow-field uncertainty from an environmental disturbance into optimization-related information and propagating it to both task-structure generation and evolutionary search regulation. The results indicate that the coordinated use of uncertainty-aware perception, structure-oriented learning, and adaptive evolutionary regulation can improve the energy efficiency, robustness, and feasibility of multi-AUV task allocation in dynamic flow fields. It can provide technical support for multi-AUV swarm operations in complex marine environments.
  • loading

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return