Skip to main navigation Skip to search Skip to main content

CLIP-HOI: Vision-Language Semantic Alignment for Human-Object Interaction Detection

  • Congcong Geng*
  • , Tian Wang
  • , Yutong Jiang
  • , Jian Wang
  • , Mali Xing
  • , Deyuan Liu
  • *Corresponding author for this work
  • Beihang University
  • China North Vehicle Research Institute
  • Wuhan University
  • Guangdong University of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

We propose CLIP-HOI, a vision-language enhanced framework for end-to-end Human-Object Interaction (HOI) detection. While transformer-based HOI methods have shown strong performance, they still struggle with rare and complex interactions that demand deeper semantic understanding beyond visual cues. To address this limitation, we integrate Contrastive Language-Image Pre-Training (CLIP) into the HoiTransformer, leveraging its powerful vision-language alignment capability. Specifically, CLIP-HOI adopts a dual-stream architecture: the original HoiTransformer branch encodes instance-level visual and spatial features, while a parallel CLIP branch extracts global, context-aware semantic representations from image-text alignment space. A cross-attention fusion module adaptively refines interaction features by infusing CLIP-guided semantic cues into visual reasoning. Extensive experiments on the HICO-DET benchmark demonstrate that CLIP-HOI consistently out-performs the baseline HoiTransformer. Our approach shows that integrating vision-language knowledge substantially improves the comprehension of human-object interactions within an end-to-end framework.

Original languageEnglish
Title of host publicationProceedings - 2025 China Automation Congress, CAC 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages7623-7628
Number of pages6
ISBN (Electronic)9798331589677
DOIs
StatePublished - 2025
Event2025 China Automation Congress, CAC 2025 - Harbin, China
Duration: 26 Sep 202528 Sep 2025

Publication series

NameProceedings - 2025 China Automation Congress, CAC 2025

Conference

Conference2025 China Automation Congress, CAC 2025
Country/TerritoryChina
CityHarbin
Period26/09/2528/09/25

Keywords

  • CLIP
  • Human-Object Interaction Detection
  • Semantic Alignment
  • Transformer
  • Vision-Language Integration

Fingerprint

Dive into the research topics of 'CLIP-HOI: Vision-Language Semantic Alignment for Human-Object Interaction Detection'. Together they form a unique fingerprint.

Cite this