Skip to main navigation Skip to search Skip to main content

Dynamic Prompting of Frozen Text-to-Image Diffusion Models for Panoptic Narrative Grounding

  • Hongyu Li
  • , Tianrui Hui
  • , Zihan Ding
  • , Jing Zhang
  • , Bin Ma
  • , Xiaoming Wei
  • , Jizhong Han
  • , Si Liu*
  • *Corresponding author for this work
  • Beihang University
  • Hefei University of Technology
  • Meituan
  • CAS - Institute of Information Engineering

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Panoptic narrative grounding (PNG), whose core target is fine-grained image-text alignment, requires a panoptic segmentation of referred objects given a narrative caption. Previous discriminative methods achieve only weak or coarse-grained alignment by panoptic segmentation pretraining or CLIP model adaptation. Given the recent progress of text-to-image Diffusion models, several works have shown their capability to achieve fine-grained image-text alignment through cross-attention maps and improved general segmentation performance. However, the direct use of phrase features as static prompts to apply frozen Diffusion models to the PNG task still suffers from a large task gap and insufficient vision-language interaction, yielding inferior performance. Therefore, we propose an Extractive-Injective Phrase Adapter (EIPA) bypass within the Diffusion UNet to dynamically update phrase prompts with image features and inject the multimodal cues back, which leverages the fine-grained image-text alignment capability of Diffusion models more sufficiently. In addition, we also design a Multi-Level Mutual Aggregation (MLMA) module to reciprocally fuse multi-level image and phrase features for segmentation refinement. Extensive experiments on the PNG benchmark show that our method achieves new state-of-the-art performance.

Original languageEnglish
Title of host publicationMM 2024 - Proceedings of the 32nd ACM International Conference on Multimedia
PublisherAssociation for Computing Machinery, Inc
Pages9485-9494
Number of pages10
ISBN (Electronic)9798400706868
DOIs
StatePublished - 28 Oct 2024
Event32nd ACM International Conference on Multimedia, MM 2024 - Melbourne, Australia
Duration: 28 Oct 20241 Nov 2024

Publication series

NameMM 2024 - Proceedings of the 32nd ACM International Conference on Multimedia

Conference

Conference32nd ACM International Conference on Multimedia, MM 2024
Country/TerritoryAustralia
CityMelbourne
Period28/10/241/11/24

Keywords

  • diffusion models
  • dynamic prompting
  • multi-level aggregation
  • panoptic narrative grounding
  • phrase adapter

Fingerprint

Dive into the research topics of 'Dynamic Prompting of Frozen Text-to-Image Diffusion Models for Panoptic Narrative Grounding'. Together they form a unique fingerprint.

Cite this