TY - GEN
T1 - PROMPT AS KNOWLEDGE BANK
T2 - 13th International Conference on Learning Representations, ICLR 2025
AU - Yang, Yuguang
AU - Chen, Tongfei
AU - Huang, Haoyu
AU - Yang, Linlin
AU - Xie, Chunyu
AU - Leng, Dawei
AU - Cao, Xianbin
AU - Zhang, Baochang
N1 - Publisher Copyright:
© 2025 13th International Conference on Learning Representations, ICLR 2025. All rights reserved.
PY - 2025
Y1 - 2025
N2 - Zero-shot medical detection enhances existing models without relying on annotated medical images, offering significant clinical value. By using grounded vision-language models (GLIP) with detailed disease descriptions as prompts, doctors can flexibly incorporate new disease characteristics to improve detection performance. However, current methods often oversimplify prompts as mere equivalents to disease names and lacks the ability to incorporate visual cues, leading to coarse image-description alignment. To address this, we propose StructuralGLIP, a framework that encodes prompts into a latent knowledge bank, enabling more context-aware and fine-grained alignment. By selecting and matching the most relevant features from image representations and the knowledge bank at layers, StructuralGLIP captures nuanced relationships between image patches and target descriptions. This approach also supports category-level prompts, which can remain fixed across all instances of the same category and provide more comprehensive information compared to instance-level prompts. Our experiments show that StructuralGLIP outperforms previous methods across various zero-shot and fine-tuned medical detection benchmarks. The code will be available at https://github.com/CapricornGuang/StructuralGLIP.
AB - Zero-shot medical detection enhances existing models without relying on annotated medical images, offering significant clinical value. By using grounded vision-language models (GLIP) with detailed disease descriptions as prompts, doctors can flexibly incorporate new disease characteristics to improve detection performance. However, current methods often oversimplify prompts as mere equivalents to disease names and lacks the ability to incorporate visual cues, leading to coarse image-description alignment. To address this, we propose StructuralGLIP, a framework that encodes prompts into a latent knowledge bank, enabling more context-aware and fine-grained alignment. By selecting and matching the most relevant features from image representations and the knowledge bank at layers, StructuralGLIP captures nuanced relationships between image patches and target descriptions. This approach also supports category-level prompts, which can remain fixed across all instances of the same category and provide more comprehensive information compared to instance-level prompts. Our experiments show that StructuralGLIP outperforms previous methods across various zero-shot and fine-tuned medical detection benchmarks. The code will be available at https://github.com/CapricornGuang/StructuralGLIP.
UR - https://www.scopus.com/pages/publications/105010250052
M3 - 会议稿件
AN - SCOPUS:105010250052
T3 - 13th International Conference on Learning Representations, ICLR 2025
SP - 53721
EP - 53743
BT - 13th International Conference on Learning Representations, ICLR 2025
PB - International Conference on Learning Representations, ICLR
Y2 - 24 April 2025 through 28 April 2025
ER -