| 摘要: |
| 智能驾驶技术的最新进展主要体现在环境感知层面, 其中传感器数据融合对提升系统性能至关重要. 点云数据虽能提供精确三维空间描述, 但存在无序性和稀疏性; 图像数据则分布规则且稠密, 二者融合可弥补单模态检测的不足. 然而, 现有融合算法存在语义信息有限、模态交互不足等问题, 多模态三维目标检测在高精度检测方面仍有提升空间. 针对此问题, 提出一种多传感器融合方法: 利用RGB图像深度补全生成伪点云, 与真实点云结合以识别感兴趣区域. 关键改进包括: 采用可变形注意力的多层次特征提取, 自适应扩展感受野至目标区域; 利用二维稀疏卷积对伪点云进行高效特征提取, 发挥其图像域规则分布特性; 提出双阶反馈机制, 在特征级通过多模态交叉注意力解决数据对齐问题, 在决策级采用高效融合策略, 实现多阶段交互训练. 该方法有效解决了伪点云精度受限与计算量增大的矛盾, 显著提升了特征提取效率与检测精度. 在KITTI数据集的实验结果表明, 所提方法在三维交通要素检测任务中实现了更优的性能, 充分验证了算法的有效性, 为智能驾驶环境感知中的多模态融合提供了新思路. |
| 关键词: 三维目标检测 多模态融合 注意力机制 点云处理 交通场景 |
| DOI:10.13328/j.cnki.jos.007545 |
| 分类号:TP181 |
| 基金项目:国家自然科学基金(62473307) |
|
| Multi-modal 3D Object Detection Method for Traffic Scenarios Based on Two-stage Feedback |
|
TANG Wen-Neng1, LI Yao-Chen1, GAO Sheng-Jing1, GAO Cong1, PENG Yue-Han1, LIU Yue-Hu2
|
|
1.School of Software Engineering, Xi’an Jiaotong University, Xi’an 710049, China;2.State Key Laboratory of Human-machine Hybrid Augmented Intelligence (Xi’an Jiaotong University), Xi’an 710049, China
|
| Abstract: |
| The latest advancements in intelligent driving technology are primarily reflected in the environmental perception layer, where sensor data fusion is critical for enhancing system performance. Although point cloud data provides accurate 3D spatial descriptions, it suffers from unorderedness and sparsity. Image data, with its regular and dense distribution, can compensate for the limitations of single-modality detection when fused with point clouds. However, existing fusion algorithms face challenges such as limited semantic information and insufficient modal interaction, leaving room for improvement in high-precision multi-modal 3D object detection. To address this issue, this study proposes an innovative multi-sensor fusion method: generating pseudo-point clouds via depth completion from RGB images and combining them with real point clouds to identify regions of interest. It introduces three key improvements: (1) deformable attention-based multi-layer feature extraction that adaptively expands the receptive field to target regions; (2) 2D sparse convolution for efficient pseudo-point cloud feature extraction leveraging their regular distribution in the image domain; and (3) a two-stage feedback mechanism employing multi-modal cross-attention at the feature level to solve data alignment issues and an efficient fusion strategy at the decision level for interactive training across different stages. These innovations effectively resolve the trade-off between pseudo-point cloud accuracy and computational load while significantly enhancing both feature extraction efficiency and detection accuracy. Experimental results on the KITTI dataset demonstrate the superior performance of the proposed method in 3D traffic object detection, validating its effectiveness and offering a new approach for multi-modal fusion in autonomous driving environmental perception. |
| Key words: 3D object detection multi-modal fusion attention mechanism point cloud processing traffic scenario |