Abstract
Multimodal large language models (MLLMs) have demonstrated strong performance in patch-level pathological image analysis; however, they often lack the holistic perceptual capability necessary for comprehensive Whole Slide Image (WSI) interpretation. Recent approaches have explored constructing slide-level MLLMs using VQA datasets that are entirely generated from pathology reports by large language models (LLMs). However, these datasets suffer from critical limitations: hallucinated content, information leakage in question stems, clinically irrelevant or visual independent questions, and the omission of essential diagnostic features-issues that undermine both data quality and clinical validity. In this paper, we introduce a clinical diagnosis template-based pipeline to collect pathological information. In collaboration with pathologists and guided by the the College of American Pathologists (CAP) Cancer Protocols, we design a Clinical Pathology Report Template (CPRT) that ensures comprehensive and standardized extraction of diagnostic elements from pathology reports. We validate the effectiveness of our pipeline on TCGA-BRCA. First, we extract pathological features from reports using CPRT. These features are then used to build CTIS-Align, a dataset of 80k slide-description pairs from 804 WSIs for vision-language alignment training, and CTISBench, a rigorously curated VQA benchmark comprising 977 WSIs and 14,879 question-answer pairs. CTIS-Bench emphasizes clinically grounded, closed-ended questions (e.g., tumor grade, receptor status) that reflect real diagnostic workflows, minimize non-visual reasoning, and require genuine slide understanding. We further propose CTIS-QA, a Slide-level Question Answering model, featuring a dual-stream architecture that mimics pathologists' diagnostic approach. One stream captures global slidelevel context via clustering-based feature aggregation, while the other focuses on salient local regions through attention-guided patch perception module. Extensive experiments on WSI-VQA, CTIS-Bench, and slide-level diagnostic tasks show that CTIS-QA consistently outperforms existing state-of-the-art models across multiple metrics. We will fully release both CTIS-Bench and CTIS-QA as open-source resources.
| Original language | English |
|---|---|
| Title of host publication | Proceedings - 2025 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2025 |
| Editors | Juan Liu, Jingshan Huang, Xiaowo Wang, Fa Zhang, Xiufen Zou, Tian Tian, Xiaohua Hu, Bin Hu, Yi Xiong |
| Publisher | Institute of Electrical and Electronics Engineers Inc. |
| Pages | 2602-2608 |
| Number of pages | 7 |
| ISBN (Electronic) | 9798331515577 |
| DOIs | |
| State | Published - 2025 |
| Event | 2025 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2025 - Wuhan, China Duration: 15 Dec 2025 → 18 Dec 2025 |
Publication series
| Name | Proceedings - 2025 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2025 |
|---|
Conference
| Conference | 2025 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2025 |
|---|---|
| Country/Territory | China |
| City | Wuhan |
| Period | 15/12/25 → 18/12/25 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 3 Good Health and Well-being
Keywords
- Computational Pathology
- Multimodal Large Language Model
- Pathology Report
- Whole Slide Image
Fingerprint
Dive into the research topics of 'CTIS-QA: Clinical Template-Informed Slide-Level Question Answering for Pathology'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver