跳到主要导航 跳到搜索 跳到主要内容

What I see is what you see: Joint attention learning for first and third person video co-analysis

  • Huangyue Yu
  • , Minjie Cai
  • , Yunfei Liu
  • , Feng Lu*
  • *此作品的通讯作者
  • Beihang University
  • Hunan University
  • Peng Cheng Laboratory

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

In recent years, more and more videos are captured from the first-person viewpoint by wearable cameras. Such first-person video provides additional information besides the traditional third-person video, and thus has a wide range of applications. However, techniques for analyzing the first-person video can be fundamentally different from those for the third-person video, and it is even more difficult to explore the shared information from both viewpoints. In this paper, we propose a novel method for first- and third-person video co-analysis. At the core of our method is the notion of “joint attention”, indicating the learnable representation that corresponds to the shared attention regions in different viewpoints and thus links the two viewpoints. To this end, we develop a multi-branch deep network with a triplet loss to extract the joint attention from the first- and third-person videos via self-supervised learning. We evaluate our method on the public dataset with cross-viewpoint video matching tasks. Our method outperforms the state-of-the-art both qualitatively and quantitatively. We also demonstrate how the learned joint attention can benefit various applications through a set of additional experiments.

源语言英语
主期刊名MM 2019 - Proceedings of the 27th ACM International Conference on Multimedia
出版商Association for Computing Machinery, Inc
1358-1366
页数9
ISBN(电子版)9781450368896
DOI
出版状态已出版 - 15 10月 2019
活动27th ACM International Conference on Multimedia, MM 2019 - Nice, 法国
期限: 21 10月 201925 10月 2019

出版系列

姓名MM 2019 - Proceedings of the 27th ACM International Conference on Multimedia

会议

会议27th ACM International Conference on Multimedia, MM 2019
国家/地区法国
Nice
时期21/10/1925/10/19

学术指纹

探究 'What I see is what you see: Joint attention learning for first and third person video co-analysis' 的科研主题。它们共同构成独一无二的学术指纹。

引用此