跳到主要导航 跳到搜索 跳到主要内容

Mixture of Attention Heads: Selecting Attention Heads Per Token

  • Xiaofeng Zhang
  • , Yikang Shen
  • , Zeyu Huang
  • , Jie Zhou
  • , Wenge Rong
  • , Zhang Xiong
  • Beihang University
  • Montreal Institute for Learning Algorithms
  • WeChat International Pte. Ltd.

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Mixture-of-Experts (MoE) networks have been proposed as an efficient way to scale up model capacity and implement conditional computing. However, the study of MoE components mostly focused on the feedforward layer in Transformer architecture. This paper proposes the Mixture of Attention Heads (MoA), a new architecture that combines multi-head attention with the MoE mechanism. MoA includes a set of attention heads that each has its own set of parameters. Given an input, a router dynamically selects a subset of k attention heads per token. This conditional computation schema allows MoA to achieve stronger performance than the standard multi-head attention layer. Furthermore, the sparsely gated MoA can easily scale up the number of attention heads and the number of parameters while preserving computational efficiency. In addition to the performance improvements, MoA also automatically differentiates heads' utilities, providing a new perspective to discuss the model's interpretability. We conducted experiments on several important tasks, including Machine Translation and Masked Language Modeling. Experiments have shown promising results on several tasks against strong baselines that involve large and very deep models.

源语言英语
主期刊名Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022
编辑Yoav Goldberg, Zornitsa Kozareva, Yue Zhang
出版商Association for Computational Linguistics (ACL)
4150-4162
页数13
ISBN(电子版)9781959429401
DOI
出版状态已出版 - 2022
活动2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022 - Hybrid, Abu Dhabi, 阿拉伯联合酋长国
期限: 7 12月 202211 12月 2022

出版系列

姓名Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022

会议

会议2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022
国家/地区阿拉伯联合酋长国
Hybrid, Abu Dhabi
时期7/12/2211/12/22

学术指纹

探究 'Mixture of Attention Heads: Selecting Attention Heads Per Token' 的科研主题。它们共同构成独一无二的学术指纹。

引用此