TY - JOUR
T1 - UnsOcc
T2 - An Occupancy Grid-Based General Detection Network for Unstructured Scenes
AU - Wang, Jie
AU - Wang, Zhangyu
AU - Yang, Songyue
AU - Luo, Wenwen
AU - Chen, Peng
AU - Yu, Guizhen
N1 - Publisher Copyright:
© 2026 IEEE. All rights reserved.
PY - 2026
Y1 - 2026
N2 - A perception and understanding of dynamic situations is a critical task for autonomous driving, particularly in complex, unstructured environments such as open-pit mining sites, where targets exhibit a vast diversity of shapes and sizes. This article introduces UnsOcc, a novel occupancy grid network specifically designed to handle such challenging scenes. The method utilizes camera images and LiDAR point clouds as inputs. A dimension lifting module (DLM) first unifies these multimodal data into a consistent voxel space through depth uncertainty estimation and cross-attention mechanisms. Subsequently, a semantic assignment module (SAM) integrates contextual and semantic features from the images with the voxel queries, effectively aligning 3-D occupancy states with their semantic information. Furthermore, a local super-resolution (LSR) module is designed to enhance the resolution of highly attended voxels, while processing less critical areas more efficiently, thus striking an optimal balance between computational cost and detection accuracy. Validated on an unstructured site dataset, UnsOcc achieves a state-of-the-art performance of 61.64% mean intersection over union (mIoU), representing a significant 6.34% improvement over baseline methods, thereby providing an effective solution for detecting and categorizing diverse objects in unstructured sites. In addition, the network has been engineered for deployment, effectively enhancing the general obstacle detection capability of mining trucks.
AB - A perception and understanding of dynamic situations is a critical task for autonomous driving, particularly in complex, unstructured environments such as open-pit mining sites, where targets exhibit a vast diversity of shapes and sizes. This article introduces UnsOcc, a novel occupancy grid network specifically designed to handle such challenging scenes. The method utilizes camera images and LiDAR point clouds as inputs. A dimension lifting module (DLM) first unifies these multimodal data into a consistent voxel space through depth uncertainty estimation and cross-attention mechanisms. Subsequently, a semantic assignment module (SAM) integrates contextual and semantic features from the images with the voxel queries, effectively aligning 3-D occupancy states with their semantic information. Furthermore, a local super-resolution (LSR) module is designed to enhance the resolution of highly attended voxels, while processing less critical areas more efficiently, thus striking an optimal balance between computational cost and detection accuracy. Validated on an unstructured site dataset, UnsOcc achieves a state-of-the-art performance of 61.64% mean intersection over union (mIoU), representing a significant 6.34% improvement over baseline methods, thereby providing an effective solution for detecting and categorizing diverse objects in unstructured sites. In addition, the network has been engineered for deployment, effectively enhancing the general obstacle detection capability of mining trucks.
KW - Instrumentation and measurement (I&M)
KW - multimodal perception
KW - semantic segmentation
KW - unstructured scenes
KW - visual-based measurement (VBM)
UR - https://www.scopus.com/pages/publications/105034677976
U2 - 10.1109/TIM.2026.3676098
DO - 10.1109/TIM.2026.3676098
M3 - 文章
AN - SCOPUS:105034677976
SN - 0018-9456
VL - 75
JO - IEEE Transactions on Instrumentation and Measurement
JF - IEEE Transactions on Instrumentation and Measurement
M1 - 5007115
ER -