Page 40 - 《软件学报》2026年第5期
P. 40
软件学报 ISSN 1000-9825, CODEN RUXUEW E-mail: jos@iscas.ac.cn
2026,37(5):1919−1935 [doi: 10.13328/j.cnki.jos.007545] [CSTR: 32375.14.jos.007545] http://www.jos.org.cn
©中国科学院软件研究所版权所有. Tel: +86-10-62562563
*
交通场景多模态双阶反馈的三维目标检测方法
唐文能 1 , 李垚辰 1 , 高笙景 1 , 高 聪 1 , 彭越涵 1 , 刘跃虎 2
1
(西安交通大学 软件学院, 陕西 西安 710049)
2
(人机混合增强智能全国重点实验室 (西安交通大学), 陕西 西安 710049)
通信作者: 李垚辰, E-mail: yaochenli@mail.xjtu.edu.cn
摘 要: 智能驾驶技术的最新进展主要体现在环境感知层面, 其中传感器数据融合对提升系统性能至关重要. 点
云数据虽能提供精确三维空间描述, 但存在无序性和稀疏性; 图像数据则分布规则且稠密, 二者融合可弥补单模
态检测的不足. 然而, 现有融合算法存在语义信息有限、模态交互不足等问题, 多模态三维目标检测在高精度检
测方面仍有提升空间. 针对此问题, 提出一种多传感器融合方法: 利用 RGB 图像深度补全生成伪点云, 与真实点
云结合以识别感兴趣区域. 关键改进包括: 采用可变形注意力的多层次特征提取, 自适应扩展感受野至目标区域;
利用二维稀疏卷积对伪点云进行高效特征提取, 发挥其图像域规则分布特性; 提出双阶反馈机制, 在特征级通过
多模态交叉注意力解决数据对齐问题, 在决策级采用高效融合策略, 实现多阶段交互训练. 该方法有效解决了伪
点云精度受限与计算量增大的矛盾, 显著提升了特征提取效率与检测精度. 在 KITTI 数据集的实验结果表明, 所
提方法在三维交通要素检测任务中实现了更优的性能, 充分验证了算法的有效性, 为智能驾驶环境感知中的多
模态融合提供了新思路.
关键词: 三维目标检测; 多模态融合; 注意力机制; 点云处理; 交通场景
中图法分类号: TP181
中文引用格式: 唐文能, 李垚辰, 高笙景, 高聪, 彭越涵, 刘跃虎. 交通场景多模态双阶反馈的三维目标检测方法. 软件学报, 2026,
37(5): 1919–1935. http://www.jos.org.cn/1000-9825/7545.htm
英文引用格式: Tang WN, Li YC, Gao SJ, Gao C, Peng YH, Liu YH. Multi-modal 3D Object Detection Method for Traffic Scenarios
Based on Two-stage Feedback. Ruan Jian Xue Bao/Journal of Software, 2026, 37(5): 1919–1935 (in Chinese). http://www.jos.org.cn/
1000-9825/7545.htm
Multi-modal 3D Object Detection Method for Traffic Scenarios Based on Two-stage Feedback
1
1
1
1
1
TANG Wen-Neng , LI Yao-Chen , GAO Sheng-Jing , GAO Cong , PENG Yue-Han , LIU Yue-Hu 2
1
(School of Software Engineering, Xi’an Jiaotong University, Xi’an 710049, China)
2
(State Key Laboratory of Human-machine Hybrid Augmented Intelligence (Xi’an Jiaotong University), Xi’an 710049, China)
Abstract: The latest advancements in intelligent driving technology are primarily reflected in the environmental perception layer, where
sensor data fusion is critical for enhancing system performance. Although point cloud data provides accurate 3D spatial descriptions, it
suffers from unorderedness and sparsity. Image data, with its regular and dense distribution, can compensate for the limitations of single-
modality detection when fused with point clouds. However, existing fusion algorithms face challenges such as limited semantic information
and insufficient modal interaction, leaving room for improvement in high-precision multi-modal 3D object detection. To address this issue,
this study proposes an innovative multi-sensor fusion method: generating pseudo-point clouds via depth completion from RGB images and
combining them with real point clouds to identify regions of interest. It introduces three key improvements: (1) deformable attention-based
* 基金项目: 国家自然科学基金 (62473307)
本文由“多媒体智能理解与生成”专题特约编辑孙立峰教授、闵巍庆副研究员、马占宇教授、蒋树强研究员、彭宇新教授、田丰研究
员、黄庆明教授推荐.
收稿时间: 2025-05-26; 修改时间: 2025-07-11; 采用时间: 2025-09-05; jos 在线出版时间: 2025-09-23
CNKI 网络首发时间: 2026-01-02

