Skip to main navigation Skip to search Skip to main content

CTIS-QA: Clinical Template-Informed Slide-Level Question Answering for Pathology

  • Hao Lu
  • , Ziniu Qian
  • , Yifu Li
  • , Yang Zhou
  • , Bingzheng Wei
  • , Yan Xu*
  • *Corresponding author for this work
  • Beihang University
  • ByteDance Ltd.

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Multimodal large language models (MLLMs) have demonstrated strong performance in patch-level pathological image analysis; however, they often lack the holistic perceptual capability necessary for comprehensive Whole Slide Image (WSI) interpretation. Recent approaches have explored constructing slide-level MLLMs using VQA datasets that are entirely generated from pathology reports by large language models (LLMs). However, these datasets suffer from critical limitations: hallucinated content, information leakage in question stems, clinically irrelevant or visual independent questions, and the omission of essential diagnostic features-issues that undermine both data quality and clinical validity. In this paper, we introduce a clinical diagnosis template-based pipeline to collect pathological information. In collaboration with pathologists and guided by the the College of American Pathologists (CAP) Cancer Protocols, we design a Clinical Pathology Report Template (CPRT) that ensures comprehensive and standardized extraction of diagnostic elements from pathology reports. We validate the effectiveness of our pipeline on TCGA-BRCA. First, we extract pathological features from reports using CPRT. These features are then used to build CTIS-Align, a dataset of 80k slide-description pairs from 804 WSIs for vision-language alignment training, and CTISBench, a rigorously curated VQA benchmark comprising 977 WSIs and 14,879 question-answer pairs. CTIS-Bench emphasizes clinically grounded, closed-ended questions (e.g., tumor grade, receptor status) that reflect real diagnostic workflows, minimize non-visual reasoning, and require genuine slide understanding. We further propose CTIS-QA, a Slide-level Question Answering model, featuring a dual-stream architecture that mimics pathologists' diagnostic approach. One stream captures global slidelevel context via clustering-based feature aggregation, while the other focuses on salient local regions through attention-guided patch perception module. Extensive experiments on WSI-VQA, CTIS-Bench, and slide-level diagnostic tasks show that CTIS-QA consistently outperforms existing state-of-the-art models across multiple metrics. We will fully release both CTIS-Bench and CTIS-QA as open-source resources.

Original languageEnglish
Title of host publicationProceedings - 2025 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2025
EditorsJuan Liu, Jingshan Huang, Xiaowo Wang, Fa Zhang, Xiufen Zou, Tian Tian, Xiaohua Hu, Bin Hu, Yi Xiong
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages2602-2608
Number of pages7
ISBN (Electronic)9798331515577
DOIs
StatePublished - 2025
Event2025 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2025 - Wuhan, China
Duration: 15 Dec 202518 Dec 2025

Publication series

NameProceedings - 2025 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2025

Conference

Conference2025 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2025
Country/TerritoryChina
CityWuhan
Period15/12/2518/12/25

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 3 - Good Health and Well-being
    SDG 3 Good Health and Well-being

Keywords

  • Computational Pathology
  • Multimodal Large Language Model
  • Pathology Report
  • Whole Slide Image

Fingerprint

Dive into the research topics of 'CTIS-QA: Clinical Template-Informed Slide-Level Question Answering for Pathology'. Together they form a unique fingerprint.

Cite this