Abstract
Pathological image analysis is a crucial field in computer-aided diagnosis. Transfer learning using models initialized on natural images has improved the downstream pathological performance. However, the lack of sophisticated domain-specific pathological initialization hinders their potential. Self-supervised learning (SSL) enables pre-training without sample-level labels, overcoming the challenge of expensive annotations. Thus, this field calls for a comprehensive dataset, similar to the ImageNet in computer vision. This work introduces a large-scale comprehensive pathological image analysis (CPIA) dataset for SSL pre-training. The CPIA dataset contains 148,962,586 images, covering over 48 organs/tissues and approximately 100 kinds of diseases, which includes two main data types: whole slide images (WSIs) and regions of interest (ROIs) images. Furthermore, we establish a standard multi-scale pathological data processing workflow, combined with the diagnosis habits of senior pathologists. The CPIA dataset facilitates a comprehensive pathological understanding and enables pattern discovery explorations. Additionally, to launch the CPIA dataset, several state-of-the-art (SOTA) baselines of SSL pre-training and downstream evaluation are specially conducted. The CPIA dataset information and code are available at https://github.com/zhanglab2021/CPIA_Dataset.
| Original language | English |
|---|---|
| Article number | 108148 |
| Journal | Biomedical Signal Processing and Control |
| Volume | 110 |
| DOIs | |
| State | Published - Dec 2025 |
Keywords
- Large-scale dataset
- Pathological images
- Pre-training
- Self-supervised learning
Fingerprint
Dive into the research topics of 'CPIA dataset: a large-scale comprehensive pathological image analysis dataset for self-supervised learning pre-training'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver