Cross-modal 3D ship object detection based on visible-light images and LiDAR point clouds
-
Abstract
Objectives To meet the high-precision sensing requirements of intelligent ships at sea, this study proposes a cross-modal 3D ship target detection method based on visible-light images and LiDAR point clouds. Methods First, a small-object detection branch, a BiFPN bidirectional feature fusion architecture, and a RepGhost lightweight module were introduced into YOLO11 to construct the RepGhost-BiFPN YOLO (RB-YOLO) 2D object detection model. This model identifies 2D candidate regions for vessels. Based on the joint camera-LiDAR calibration relationship, the corresponding local view cone point clouds are extracted, transforming the global 3D object search into local 3D bounding box regression within the candidate regions; Second, a feature encoding module sensitive to the reflectance intensity of the LiDAR point cloud is constructed. This module combines LiDAR reflectance intensity with 3D coordinate features and utilizes a multi-scale set of point cloud features along with a Transformer encoder to extract local geometric features of the vessel; finally, a cross-modal attention mechanism is introduced to aggregate texture, contour, and semantic information from the image, enabling the detection of the vessel’s 3D center, scale, and heading. Results Experimental results show that, in the evaluation of the 3D detection model, the proposed method achieves AP3D@0.25, AP3D@0.50, and AP3D@0.70 values of 90.60%, 79.78%, and 58.72%, respectively. During the departure test of the smart vessel “Xin Hongzhuan,” the mean absolute errors of target-vessel distance and heading were 1.29 m and 3.02°, respectively. Conclusions This study validates the effectiveness of the cross-modal 3D vessel detection method based on visible-light images and LiDAR point clouds, providing a new technical approach for acquiring 3D spatial information of vessel targets in complex maritime environments.
-
-