TY - GEN
T1 - FCMF-UNetFormer
T2 - 2025 5th International Conference on Communication Technology and Information Technology, ICCTIT 2025
AU - Liu, Zhuo
AU - Jia, Yanxi
AU - Zhang, Chuang
AU - Cheng, Jingchun
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - With the rapid growth of high-resolution remote sensing imagery, fine-grained urban scene parsing faces increasing challenges such as diverse object categories, large scale variations, and complex boundaries. Towards this issue, traditional CNNs are limited by local receptive fields; while Transformers which are, strong in global modeling struggle with hierarchical feature fusion. In this paper, we propose the FCMF-UNetFormer for urban scene parsing, a fully connected multi-scale fusion network which solves the semantic gap between shallow and deep features of transformer. In specific, we introduce a fully connected skip structure to enable progressive multi-scale interaction, a MultiSkipFusion module for continuous deep-shallow semantic integration, and a Proportional Skip Fusion (PSF) module that adaptively aligns and fuses features through learnable proportional weights. Experiments on the Vaihingen and LoveDA datasets demonstrate that the FCMF-UNetFormer achieves notable improvements over its baseline (i.e. UNetFormer) and outperforms recent state-of-the-art models in mIoU, F1-score, boundary quality, and overall accuracy, showing strong robustness and generalization.
AB - With the rapid growth of high-resolution remote sensing imagery, fine-grained urban scene parsing faces increasing challenges such as diverse object categories, large scale variations, and complex boundaries. Towards this issue, traditional CNNs are limited by local receptive fields; while Transformers which are, strong in global modeling struggle with hierarchical feature fusion. In this paper, we propose the FCMF-UNetFormer for urban scene parsing, a fully connected multi-scale fusion network which solves the semantic gap between shallow and deep features of transformer. In specific, we introduce a fully connected skip structure to enable progressive multi-scale interaction, a MultiSkipFusion module for continuous deep-shallow semantic integration, and a Proportional Skip Fusion (PSF) module that adaptively aligns and fuses features through learnable proportional weights. Experiments on the Vaihingen and LoveDA datasets demonstrate that the FCMF-UNetFormer achieves notable improvements over its baseline (i.e. UNetFormer) and outperforms recent state-of-the-art models in mIoU, F1-score, boundary quality, and overall accuracy, showing strong robustness and generalization.
KW - FCMF-UNetFormer
KW - Multi-scale fusion
KW - Remote sensing
KW - Semantic segmentation
KW - Transformer
KW - UNetFormer
KW - Urban scenes parsing
UR - https://www.scopus.com/pages/publications/105036150122
U2 - 10.1109/ICCTIT68197.2025.11406413
DO - 10.1109/ICCTIT68197.2025.11406413
M3 - 会议稿件
AN - SCOPUS:105036150122
T3 - 2025 5th International Conference on Communication Technology and Information Technology, ICCTIT 2025
SP - 232
EP - 237
BT - 2025 5th International Conference on Communication Technology and Information Technology, ICCTIT 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 26 December 2025 through 28 December 2025
ER -