TY - JOUR
T1 - HiViT
T2 - Hierarchical attention-based Transformer for multi-scale whole slide histopathological image classification
AU - Yu, Jinze
AU - Li, Shuo
AU - Tan, Luxin
AU - Zhou, Haoyi
AU - Li, Zhongwu
AU - Li, Jianxin
N1 - Publisher Copyright:
© 2025 Elsevier Ltd
PY - 2025/6/5
Y1 - 2025/6/5
N2 - The classification of Whole Slide Images (WSIs) remains a challenging task in computer-aided diagnostics. Though Multi-Instance Learning (MIL) and patch-wise modeling have become the mainstream methods in current research, they rely on a strong assumption that different patches are independent and identically distributed. Therefore, the contextual correlations within the images and across all the patches have been neglected, resulting in inferior performance, particularly in global-level tasks like mutation predictions. In this paper, we propose HiViT, a multi-scale WSI classification Transformer model utilizing the proposed Cross-scale Hierarchical Self-Attention (CHSA) mechanism to aggregate patch features among different scales and capture context information from the entire image with a reasonable memory cost. The CHSA mechanism is developed based on the multiple magnification scannings of WSIs and restricts cross-scale self-attention computation within local-constrained windows to aggregate patch features among different scales and capture context information from the entire image while getting rid of the infeasible memory cost of vanilla Transformers, and finally achieves the first fully trainable Transformer model for the WSI classification problem. Additionally, the Multi-Scale Feature Augmentation method is proposed to enhance permutation invariance on a large scale, which is important in the MIL assumption, while preserving local fine-grained contextual correlations. We extensively evaluated HiViT on several real-world datasets of diverse tasks, and our model outperformed the current state-of-the-art Transformer and MIL-based methods. Code will be available at https://github.com/BUAA-SMART-Med-CV/HiViT.
AB - The classification of Whole Slide Images (WSIs) remains a challenging task in computer-aided diagnostics. Though Multi-Instance Learning (MIL) and patch-wise modeling have become the mainstream methods in current research, they rely on a strong assumption that different patches are independent and identically distributed. Therefore, the contextual correlations within the images and across all the patches have been neglected, resulting in inferior performance, particularly in global-level tasks like mutation predictions. In this paper, we propose HiViT, a multi-scale WSI classification Transformer model utilizing the proposed Cross-scale Hierarchical Self-Attention (CHSA) mechanism to aggregate patch features among different scales and capture context information from the entire image with a reasonable memory cost. The CHSA mechanism is developed based on the multiple magnification scannings of WSIs and restricts cross-scale self-attention computation within local-constrained windows to aggregate patch features among different scales and capture context information from the entire image while getting rid of the infeasible memory cost of vanilla Transformers, and finally achieves the first fully trainable Transformer model for the WSI classification problem. Additionally, the Multi-Scale Feature Augmentation method is proposed to enhance permutation invariance on a large scale, which is important in the MIL assumption, while preserving local fine-grained contextual correlations. We extensively evaluated HiViT on several real-world datasets of diverse tasks, and our model outperformed the current state-of-the-art Transformer and MIL-based methods. Code will be available at https://github.com/BUAA-SMART-Med-CV/HiViT.
KW - Deep learning
KW - Histopathological image classification
KW - Medical image analysis
KW - Vision Transformer
UR - https://www.scopus.com/pages/publications/105000121100
U2 - 10.1016/j.eswa.2025.127164
DO - 10.1016/j.eswa.2025.127164
M3 - 文章
AN - SCOPUS:105000121100
SN - 0957-4174
VL - 277
JO - Expert Systems with Applications
JF - Expert Systems with Applications
M1 - 127164
ER -