Skip to main navigation Skip to search Skip to main content

PROMPT AS KNOWLEDGE BANK: BOOST VISION-LANGUAGE MODEL VIA STRUCTURAL REPRESENTATION FOR ZERO-SHOT MEDICAL DETECTION

  • Yuguang Yang
  • , Tongfei Chen
  • , Haoyu Huang
  • , Linlin Yang*
  • , Chunyu Xie*
  • , Dawei Leng
  • , Xianbin Cao
  • , Baochang Zhang
  • *Corresponding author for this work
  • Beihang University
  • Qihoo 360
  • Communication University of China
  • Lobachevsky State University of Nizhni Novgorod

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Zero-shot medical detection enhances existing models without relying on annotated medical images, offering significant clinical value. By using grounded vision-language models (GLIP) with detailed disease descriptions as prompts, doctors can flexibly incorporate new disease characteristics to improve detection performance. However, current methods often oversimplify prompts as mere equivalents to disease names and lacks the ability to incorporate visual cues, leading to coarse image-description alignment. To address this, we propose StructuralGLIP, a framework that encodes prompts into a latent knowledge bank, enabling more context-aware and fine-grained alignment. By selecting and matching the most relevant features from image representations and the knowledge bank at layers, StructuralGLIP captures nuanced relationships between image patches and target descriptions. This approach also supports category-level prompts, which can remain fixed across all instances of the same category and provide more comprehensive information compared to instance-level prompts. Our experiments show that StructuralGLIP outperforms previous methods across various zero-shot and fine-tuned medical detection benchmarks. The code will be available at https://github.com/CapricornGuang/StructuralGLIP.

Original languageEnglish
Title of host publication13th International Conference on Learning Representations, ICLR 2025
PublisherInternational Conference on Learning Representations, ICLR
Pages53721-53743
Number of pages23
ISBN (Electronic)9798331320850
StatePublished - 2025
Event13th International Conference on Learning Representations, ICLR 2025 - Singapore, Singapore
Duration: 24 Apr 202528 Apr 2025

Publication series

Name13th International Conference on Learning Representations, ICLR 2025

Conference

Conference13th International Conference on Learning Representations, ICLR 2025
Country/TerritorySingapore
CitySingapore
Period24/04/2528/04/25

Fingerprint

Dive into the research topics of 'PROMPT AS KNOWLEDGE BANK: BOOST VISION-LANGUAGE MODEL VIA STRUCTURAL REPRESENTATION FOR ZERO-SHOT MEDICAL DETECTION'. Together they form a unique fingerprint.

Cite this