Normalizing flow-based action-constrained TD3 control method for underwater dynamic docking
-
Abstract
Objectives In underwater dynamic docking processes, due to the complex flow field disturbances induced by relative variable-speed motion and the docking mechanism, thruster inputs are more likely to trigger physical constraints such as saturation and dead zones, making it difficult for traditional reinforcement learning methods to achieve stable and efficient learning while guaranteeing action feasibility. To address this problem, this paper proposes a normalizing flow–based action-constrained Twin Delayed Deep Deterministic Policy Gradient (TD3) control method for underwater dynamic docking. Methods Within the TD3 framework, the proposed method introduces a normalizing flow to explicitly model the feasible action space of the thrusters. By constructing an invertible and differentiable mapping from latent actions to feasible actions, it embeds action constraints directly into the policy generation process, thereby improving control-action feasibility and reducing the frequency of constraint violations. Based on a conditional affine normalizing flow structure, the explicit mathematical formulation of the action mapping is provided, and the deterministic policy gradient with the incorporation of the normalizing flow is derived, avoiding the zero-gradient issue commonly encountered in traditional action projection methods. On this basis, an action-constrained TD3 algorithm based on normalizing flows is constructed, and the complete algorithmic procedure is presented. Results Comparative numerical simulations show that the proposed method improves control accuracy and convergence stability in underwater dynamic docking while maintaining control-input feasibility. Compared with TD3 with action projection, the terminal position error is reduced by 47.0%, while the numbers of constraint-boundary contacts in the yaw-moment and surge-force channels are reduced by 66.4% and 89.5%, respectively. Tank experiments on the heading control of a catamaran unmanned surface vehicle further show that the proposed method reduces the heading-control root-mean-square error by 26.0% and the number of action-boundary contacts of the normalized yaw control input by 95.7%. These results verify the real-time deployability and action-constraint handling capability of the proposed method in the yaw channel. Conclusions The proposed method can achieve underwater dynamic docking control while satisfying the physical constraints of thrusters, improving docking accuracy and convergence stability. It shows potential advantages in ensuring action feasibility under complex flow-field disturbances and enhancing the control performance of systems with constrained actuators.
-
-