跳到主要导航 跳到搜索 跳到主要内容

MV-ViT: High-Quality Features and Simple Fusion for Superior Multi-View 3D Object Classification

  • Beihang University
  • Commercial Aircraft Corporation of China, Ltd.

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

We present MV-ViT, a novel and efficient approach for multi-view 3D object classification. MV-ViT leverages the power of pre-trained Vision Transformers (ViTs) for robust single-view feature extraction, coupled with a two-stage training strategy that enhances performance while minimizing computational overhead. Unlike methods relying on complex multi-view fusion mechanisms, we demonstrate that high-quality features extracted by a fine-tuned ViT, when aggregated via simple mean pooling, achieve state-of-the-art results. MV-ViT achieves 97.23% accuracy with 12 views on the ModelNet40 benchmark, surpassing methods employing significantly more elaborate fusion architectures. Ablation studies highlight the critical role of strong ViT-derived features in achieving superior performance, making complex fusion redundant.

源语言英语
主期刊名Proceedings - 2025 China Automation Congress, CAC 2025
出版商Institute of Electrical and Electronics Engineers Inc.
120-125
页数6
ISBN(电子版)9798331589677
DOI
出版状态已出版 - 2025
活动2025 China Automation Congress, CAC 2025 - Harbin, 中国
期限: 26 9月 202528 9月 2025

出版系列

姓名Proceedings - 2025 China Automation Congress, CAC 2025

会议

会议2025 China Automation Congress, CAC 2025
国家/地区中国
Harbin
时期26/09/2528/09/25

学术指纹

探究 'MV-ViT: High-Quality Features and Simple Fusion for Superior Multi-View 3D Object Classification' 的科研主题。它们共同构成独一无二的学术指纹。

引用此