TY - JOUR
T1 - CPIA dataset
T2 - a large-scale comprehensive pathological image analysis dataset for self-supervised learning pre-training
AU - Ying, Nan
AU - Lei, Yanli
AU - Zhang, Tianyi
AU - Lyu, Shangqing
AU - Chen, Sicheng
AU - Liu, Zeyu
AU - Feng, Yunlu
AU - Zhao, Yu
AU - Zhang, Guanglei
N1 - Publisher Copyright:
© 2025 Elsevier Ltd
PY - 2025/12
Y1 - 2025/12
N2 - Pathological image analysis is a crucial field in computer-aided diagnosis. Transfer learning using models initialized on natural images has improved the downstream pathological performance. However, the lack of sophisticated domain-specific pathological initialization hinders their potential. Self-supervised learning (SSL) enables pre-training without sample-level labels, overcoming the challenge of expensive annotations. Thus, this field calls for a comprehensive dataset, similar to the ImageNet in computer vision. This work introduces a large-scale comprehensive pathological image analysis (CPIA) dataset for SSL pre-training. The CPIA dataset contains 148,962,586 images, covering over 48 organs/tissues and approximately 100 kinds of diseases, which includes two main data types: whole slide images (WSIs) and regions of interest (ROIs) images. Furthermore, we establish a standard multi-scale pathological data processing workflow, combined with the diagnosis habits of senior pathologists. The CPIA dataset facilitates a comprehensive pathological understanding and enables pattern discovery explorations. Additionally, to launch the CPIA dataset, several state-of-the-art (SOTA) baselines of SSL pre-training and downstream evaluation are specially conducted. The CPIA dataset information and code are available at https://github.com/zhanglab2021/CPIA_Dataset.
AB - Pathological image analysis is a crucial field in computer-aided diagnosis. Transfer learning using models initialized on natural images has improved the downstream pathological performance. However, the lack of sophisticated domain-specific pathological initialization hinders their potential. Self-supervised learning (SSL) enables pre-training without sample-level labels, overcoming the challenge of expensive annotations. Thus, this field calls for a comprehensive dataset, similar to the ImageNet in computer vision. This work introduces a large-scale comprehensive pathological image analysis (CPIA) dataset for SSL pre-training. The CPIA dataset contains 148,962,586 images, covering over 48 organs/tissues and approximately 100 kinds of diseases, which includes two main data types: whole slide images (WSIs) and regions of interest (ROIs) images. Furthermore, we establish a standard multi-scale pathological data processing workflow, combined with the diagnosis habits of senior pathologists. The CPIA dataset facilitates a comprehensive pathological understanding and enables pattern discovery explorations. Additionally, to launch the CPIA dataset, several state-of-the-art (SOTA) baselines of SSL pre-training and downstream evaluation are specially conducted. The CPIA dataset information and code are available at https://github.com/zhanglab2021/CPIA_Dataset.
KW - Large-scale dataset
KW - Pathological images
KW - Pre-training
KW - Self-supervised learning
UR - https://www.scopus.com/pages/publications/105006724522
U2 - 10.1016/j.bspc.2025.108148
DO - 10.1016/j.bspc.2025.108148
M3 - 文章
AN - SCOPUS:105006724522
SN - 1746-8094
VL - 110
JO - Biomedical Signal Processing and Control
JF - Biomedical Signal Processing and Control
M1 - 108148
ER -