Skip to main navigation Skip to search Skip to main content

STPNet: Scale-Aware Text Prompt Network for Medical Image Segmentation

  • Dandan Shan
  • , Zihan Li
  • , Yunxiang Li
  • , Qingde Li
  • , Jie Tian
  • , Qingqi Hong*
  • *Corresponding author for this work
  • Xiamen University
  • University of Washington
  • University of Texas Southwestern Medical Center
  • University of Hull
  • CAS - Institute of Automation

Research output: Contribution to journalArticlepeer-review

Abstract

Accurate segmentation of lesions plays a critical role in medical image analysis and diagnosis. Traditional segmentation approaches that rely solely on visual features often struggle with the inherent uncertainty in lesion distribution and size. To address these issues, we propose STPNet, a Scaleaware Text Prompt Network that leverages vision-language modeling to enhance medical image segmentation. Our approach utilizes multi-scale textual descriptions to guide lesion localization and employs retrieval-segmentation joint learning to bridge the semantic gap between visual and linguistic modalities. Crucially, STPNet retrieves relevant textual information from a specialized medical text repository during training, eliminating the need for text input during inference while retaining the benefits of cross-modal learning. We evaluate STPNet on three datasets: COVID-Xray, COVID-CT, and Kvasir-SEG. Experimental results show that our vision-language approach outperforms state-ofthe-art segmentation methods, demonstrating the effectiveness of incorporating textual semantic knowledge into medical image analysis.

Original languageEnglish
Pages (from-to)3169-3180
Number of pages12
JournalIEEE Transactions on Image Processing
Volume34
DOIs
StatePublished - 2025
Externally publishedYes

Keywords

  • Multi-scale learning
  • medical image segmentation
  • text prompt

Fingerprint

Dive into the research topics of 'STPNet: Scale-Aware Text Prompt Network for Medical Image Segmentation'. Together they form a unique fingerprint.

Cite this