跳到主要导航 跳到搜索 跳到主要内容

PROMPT AS KNOWLEDGE BANK: BOOST VISION-LANGUAGE MODEL VIA STRUCTURAL REPRESENTATION FOR ZERO-SHOT MEDICAL DETECTION

  • Yuguang Yang
  • , Tongfei Chen
  • , Haoyu Huang
  • , Linlin Yang*
  • , Chunyu Xie*
  • , Dawei Leng
  • , Xianbin Cao
  • , Baochang Zhang
  • *此作品的通讯作者
  • Beihang University
  • Qihoo 360
  • Communication University of China
  • Lobachevsky State University of Nizhni Novgorod

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Zero-shot medical detection enhances existing models without relying on annotated medical images, offering significant clinical value. By using grounded vision-language models (GLIP) with detailed disease descriptions as prompts, doctors can flexibly incorporate new disease characteristics to improve detection performance. However, current methods often oversimplify prompts as mere equivalents to disease names and lacks the ability to incorporate visual cues, leading to coarse image-description alignment. To address this, we propose StructuralGLIP, a framework that encodes prompts into a latent knowledge bank, enabling more context-aware and fine-grained alignment. By selecting and matching the most relevant features from image representations and the knowledge bank at layers, StructuralGLIP captures nuanced relationships between image patches and target descriptions. This approach also supports category-level prompts, which can remain fixed across all instances of the same category and provide more comprehensive information compared to instance-level prompts. Our experiments show that StructuralGLIP outperforms previous methods across various zero-shot and fine-tuned medical detection benchmarks. The code will be available at https://github.com/CapricornGuang/StructuralGLIP.

源语言英语
主期刊名13th International Conference on Learning Representations, ICLR 2025
出版商International Conference on Learning Representations, ICLR
53721-53743
页数23
ISBN(电子版)9798331320850
出版状态已出版 - 2025
活动13th International Conference on Learning Representations, ICLR 2025 - Singapore, 新加坡
期限: 24 4月 202528 4月 2025

出版系列

姓名13th International Conference on Learning Representations, ICLR 2025

会议

会议13th International Conference on Learning Representations, ICLR 2025
国家/地区新加坡
Singapore
时期24/04/2528/04/25

指纹

探究 'PROMPT AS KNOWLEDGE BANK: BOOST VISION-LANGUAGE MODEL VIA STRUCTURAL REPRESENTATION FOR ZERO-SHOT MEDICAL DETECTION' 的科研主题。它们共同构成独一无二的指纹。

引用此