TY - GEN
T1 - DEFORMABLE SPHERICAL GEOMETRY TRANSFORMER FOR PANORAMIC SEMANTIC SEGMENTATION
AU - Lan, Boyang
AU - Yang, Li
AU - Xu, Mai
AU - Jiang, Lai
AU - Wang, Yufeng
N1 - Publisher Copyright:
©2025 IEEE.
PY - 2025
Y1 - 2025
N2 - The increasing availability of 360◦ images has created a demand for effective Panoramic Semantic Segmentation (PASS) to enable comprehensive scene understanding. However, the spherical nature of 360◦ image introduces significant spatial distortions due to Equirectangular Projection (ERP), making it challenging for traditional 2D methods, which are designed for Euclidean spaces. Existing PASS methods typically mitigate these distortions through developing spherical-to-tangent polyhedron transformations or specialized convolutional structures. Nevertheless, these approaches still struggle to preserve the spherical geometry and fail to adequately capture the semantic context of 360◦ images. In this paper, we propose a Deformable Spherical Geometry Transformer (DSGT) network that adapts to spherical distortions through a local-global self-attention mechanism. The local self-attention module captures local semantic information to alleviate distortions, while the global self-attention module integrates spherical geometric priors to enhance predictions. Experimental results on the Stanford2D3D panoramic dataset demonstrate that DSGT outperforms state-of-the-art PASS methods.
AB - The increasing availability of 360◦ images has created a demand for effective Panoramic Semantic Segmentation (PASS) to enable comprehensive scene understanding. However, the spherical nature of 360◦ image introduces significant spatial distortions due to Equirectangular Projection (ERP), making it challenging for traditional 2D methods, which are designed for Euclidean spaces. Existing PASS methods typically mitigate these distortions through developing spherical-to-tangent polyhedron transformations or specialized convolutional structures. Nevertheless, these approaches still struggle to preserve the spherical geometry and fail to adequately capture the semantic context of 360◦ images. In this paper, we propose a Deformable Spherical Geometry Transformer (DSGT) network that adapts to spherical distortions through a local-global self-attention mechanism. The local self-attention module captures local semantic information to alleviate distortions, while the global self-attention module integrates spherical geometric priors to enhance predictions. Experimental results on the Stanford2D3D panoramic dataset demonstrate that DSGT outperforms state-of-the-art PASS methods.
KW - 360° image
KW - semantic segmentation
KW - spherical distortions
KW - transformer
UR - https://www.scopus.com/pages/publications/105028604270
U2 - 10.1109/ICIP55913.2025.11084725
DO - 10.1109/ICIP55913.2025.11084725
M3 - 会议稿件
AN - SCOPUS:105028604270
T3 - Proceedings - International Conference on Image Processing, ICIP
SP - 271
EP - 276
BT - 2025 IEEE International Conference on Image Processing, ICIP 2025 - Proceedings
PB - IEEE Computer Society
T2 - 32nd IEEE International Conference on Image Processing, ICIP 2025
Y2 - 14 September 2025 through 17 September 2025
ER -