跳到主要导航 跳到搜索 跳到主要内容

CTIS-QA: Clinical Template-Informed Slide-Level Question Answering for Pathology

  • Hao Lu
  • , Ziniu Qian
  • , Yifu Li
  • , Yang Zhou
  • , Bingzheng Wei
  • , Yan Xu*
  • *此作品的通讯作者
  • Beihang University
  • ByteDance Ltd.

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Multimodal large language models (MLLMs) have demonstrated strong performance in patch-level pathological image analysis; however, they often lack the holistic perceptual capability necessary for comprehensive Whole Slide Image (WSI) interpretation. Recent approaches have explored constructing slide-level MLLMs using VQA datasets that are entirely generated from pathology reports by large language models (LLMs). However, these datasets suffer from critical limitations: hallucinated content, information leakage in question stems, clinically irrelevant or visual independent questions, and the omission of essential diagnostic features-issues that undermine both data quality and clinical validity. In this paper, we introduce a clinical diagnosis template-based pipeline to collect pathological information. In collaboration with pathologists and guided by the the College of American Pathologists (CAP) Cancer Protocols, we design a Clinical Pathology Report Template (CPRT) that ensures comprehensive and standardized extraction of diagnostic elements from pathology reports. We validate the effectiveness of our pipeline on TCGA-BRCA. First, we extract pathological features from reports using CPRT. These features are then used to build CTIS-Align, a dataset of 80k slide-description pairs from 804 WSIs for vision-language alignment training, and CTISBench, a rigorously curated VQA benchmark comprising 977 WSIs and 14,879 question-answer pairs. CTIS-Bench emphasizes clinically grounded, closed-ended questions (e.g., tumor grade, receptor status) that reflect real diagnostic workflows, minimize non-visual reasoning, and require genuine slide understanding. We further propose CTIS-QA, a Slide-level Question Answering model, featuring a dual-stream architecture that mimics pathologists' diagnostic approach. One stream captures global slidelevel context via clustering-based feature aggregation, while the other focuses on salient local regions through attention-guided patch perception module. Extensive experiments on WSI-VQA, CTIS-Bench, and slide-level diagnostic tasks show that CTIS-QA consistently outperforms existing state-of-the-art models across multiple metrics. We will fully release both CTIS-Bench and CTIS-QA as open-source resources.

源语言英语
主期刊名Proceedings - 2025 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2025
编辑Juan Liu, Jingshan Huang, Xiaowo Wang, Fa Zhang, Xiufen Zou, Tian Tian, Xiaohua Hu, Bin Hu, Yi Xiong
出版商Institute of Electrical and Electronics Engineers Inc.
2602-2608
页数7
ISBN(电子版)9798331515577
DOI
出版状态已出版 - 2025
活动2025 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2025 - Wuhan, 中国
期限: 15 12月 202518 12月 2025

出版系列

姓名Proceedings - 2025 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2025

会议

会议2025 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2025
国家/地区中国
Wuhan
时期15/12/2518/12/25

联合国可持续发展目标

此成果有助于实现下列可持续发展目标:

  1. 可持续发展目标 3 - 良好健康与福祉
    可持续发展目标 3 良好健康与福祉

学术指纹

探究 'CTIS-QA: Clinical Template-Informed Slide-Level Question Answering for Pathology' 的科研主题。它们共同构成独一无二的学术指纹。

引用此