跳到主要导航 跳到搜索 跳到主要内容

MRCap: Multi-modal and Multi-level Relationship-based Dense Video Captioning

  • Zhongguancun Laboratory
  • Zhengzhou University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Dense video captioning, with the objective of describing a sequence of events in a video, has received much attention recently. As events in a video are highly correlated, leveraging relationships among events helps generate coherent captions. To utilize relationships among events, existing methods mainly enrich event representations with their context, either in the form of vision (i.e., video segments) or combining vision and language (i.e., captions). However, these methods do not explicitly exploit the correspondence between these two modalities. Moreover, the video-level context spanning multiple events is not fully exploited. In this paper, we propose MRCap, a novel relationship-based model for dense video captioning. The key of MRCap is a multi-modal and multi-level event relationship module (MMERM). MMERM exploits the correspondence between vision and language at both the event level and the video level via contrastive learning. Experiments on ActivityNet Captions and YouCook2 datasets demonstrate that MRCap achieves state-of-the-art performance.

源语言英语
主期刊名Proceedings - 2023 IEEE International Conference on Multimedia and Expo, ICME 2023
出版商IEEE Computer Society
2615-2620
页数6
ISBN(电子版)9781665468916
DOI
出版状态已出版 - 2023
活动2023 IEEE International Conference on Multimedia and Expo, ICME 2023 - Brisbane, 澳大利亚
期限: 10 7月 202314 7月 2023

出版系列

姓名Proceedings - IEEE International Conference on Multimedia and Expo
2023-July
ISSN(印刷版)1945-7871
ISSN(电子版)1945-788X

会议

会议2023 IEEE International Conference on Multimedia and Expo, ICME 2023
国家/地区澳大利亚
Brisbane
时期10/07/2314/07/23

指纹

探究 'MRCap: Multi-modal and Multi-level Relationship-based Dense Video Captioning' 的科研主题。它们共同构成独一无二的指纹。

引用此