跳到主要导航 跳到搜索 跳到主要内容

DiffDVC: Accurate Event Detection for Dense Video Captioning via Diffusion Models

  • Zhongguancun Laboratory
  • Zhengzhou University
  • Beihang University
  • SUNY Buffalo

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Dense video captioning (DVC) aims to describe multiple events within a video, and its performance is greatly affected by the accuracy of video event detection. Video event detection involves predicting the proposal boundaries (start and end times) and the classification score of each event in a video. Recently, a few methods have applied diffusion models originally designed for image object detection to detect events in DVC. These methods add noise to the ground-truth event proposal boundaries, and subsequently learn the denoising process. However, these methods often overlook the fundamental differences between videos and images. We observe that, whereas in images the important information for object classification is normally around the boundaries of the ground-truth boxes, in videos the key information for event classification is typically centered in the middle of ground-truth event proposals. As a result, the classification module in these existing diffusion models becomes insensitive to boundary changes introduced by the added noise, leading to suboptimal performance. This paper introduces DiffDVC, an innovative diffusion model for DVC. The core of DiffDVC is a boundary-sensitive detector. The detector increases the sensitivity of the classification module to boundary changes by focusing on frames within a specific range around the start and end times of noisy event proposals. Additionally, this range is dynamically adjusted to suit different event proposals. Comprehensive experiments on ActivityNet-1.3, ActivityNet Captions, and YouCook2 datasets show DiffDVC achieving superior performance.

源语言英语
主期刊名Special Track on AI Alignment
编辑Toby Walsh, Julie Shah, Zico Kolter
出版商Association for the Advancement of Artificial Intelligence
2221-2229
页数9
版本2
ISBN(电子版)157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 157735897X, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978, 9781577358978
DOI
出版状态已出版 - 11 4月 2025
活动39th Annual AAAI Conference on Artificial Intelligence, AAAI 2025 - Philadelphia, 美国
期限: 25 2月 20254 3月 2025

出版系列

姓名Proceedings of the AAAI Conference on Artificial Intelligence
编号2
39
ISSN(印刷版)2159-5399
ISSN(电子版)2374-3468

会议

会议39th Annual AAAI Conference on Artificial Intelligence, AAAI 2025
国家/地区美国
Philadelphia
时期25/02/254/03/25

指纹

探究 'DiffDVC: Accurate Event Detection for Dense Video Captioning via Diffusion Models' 的科研主题。它们共同构成独一无二的指纹。

引用此