TY - GEN
T1 - DSA3D
T2 - 2026 3rd International Conference on Autonomous Driving and Intelligent Sensing Technology, ADIST 2026
AU - Wang, Shuoheng
AU - Chen, Jie
AU - Wan, Huiyao
N1 - Publisher Copyright:
© 2026 Copyright held by the owner/author(s).
PY - 2026/5/13
Y1 - 2026/5/13
N2 - Accurate 3D object detection utilizing sparse LiDAR point clouds stands as a critical bottleneck in autonomous perception, largely stemming from the data's intrinsic irregularity and varying spatial density. Existing voxel-based paradigms often suffer from feature quantization loss, where heuristic pooling operations discard critical micro-geometric details, while point-based methods incur prohibitive computational costs. Toward this end, we propose DSA3D, an integrated framework that synergizes hierarchical feature abstraction with dynamic computation.Specifically, we introduce a Hierarchical Voxel Feature Encoder incorporating the Spatial-Voxel Attention (SVA) mechanism. By leveraging a dual-stage residual design, SVA explicitly models intra-voxel geometric distributions and inter-channel semantic dependencies, effectively mitigating quantization artifacts. Furthermore, we design a Spatially-Adaptive 3D Backbone equipped with Dynamic Sparse Convolution (DSC). This operator functions as a learnable importance evaluator, adaptively allocating computational budget to semantically salient regions while pruning redundant background noise. Finally, we present the Spatial Context Aggregation Network (SCA-Net) to facilitate wide-area reasoning in the BEV domain. Thorough evaluations on the KITTI benchmark reveal that DSA3D attains highly competitive performance, yielding 82.53% AP for the Car category.
AB - Accurate 3D object detection utilizing sparse LiDAR point clouds stands as a critical bottleneck in autonomous perception, largely stemming from the data's intrinsic irregularity and varying spatial density. Existing voxel-based paradigms often suffer from feature quantization loss, where heuristic pooling operations discard critical micro-geometric details, while point-based methods incur prohibitive computational costs. Toward this end, we propose DSA3D, an integrated framework that synergizes hierarchical feature abstraction with dynamic computation.Specifically, we introduce a Hierarchical Voxel Feature Encoder incorporating the Spatial-Voxel Attention (SVA) mechanism. By leveraging a dual-stage residual design, SVA explicitly models intra-voxel geometric distributions and inter-channel semantic dependencies, effectively mitigating quantization artifacts. Furthermore, we design a Spatially-Adaptive 3D Backbone equipped with Dynamic Sparse Convolution (DSC). This operator functions as a learnable importance evaluator, adaptively allocating computational budget to semantically salient regions while pruning redundant background noise. Finally, we present the Spatial Context Aggregation Network (SCA-Net) to facilitate wide-area reasoning in the BEV domain. Thorough evaluations on the KITTI benchmark reveal that DSA3D attains highly competitive performance, yielding 82.53% AP for the Car category.
KW - 3D object detection
KW - LiDAR point cloud
KW - Sparse convolution
UR - https://www.scopus.com/pages/publications/105040241123
U2 - 10.1145/3807211.3807231
DO - 10.1145/3807211.3807231
M3 - 会议稿件
AN - SCOPUS:105040241123
T3 - Proceedings of 2026 3rd International Conference on Autonomous Driving and Intelligent Sensing Technology, ADIST 2026
SP - 140
EP - 147
BT - Proceedings of 2026 3rd International Conference on Autonomous Driving and Intelligent Sensing Technology, ADIST 2026
PB - Association for Computing Machinery, Inc
Y2 - 6 February 2026 through 8 February 2026
ER -