TY - GEN
T1 - History attention for source-target alignment in neural machine translation
AU - Yu, Yuanyuan
AU - Huang, Yan
AU - Chao, Wenhan
AU - Zhang, Peidong
N1 - Publisher Copyright:
© 2018 IEEE.
PY - 2018/6/8
Y1 - 2018/6/8
N2 - Attention mechanism has enhanced state-of-the-art Neural Machine Translation (NMT) by focusing on parts of the source sentence when predicting each target word. However we find that most of the attention context vector calculation is directly dependent on the current decoder hidden state. It tends to ignore past translated information, which often leads to over-translation and under-translation. When target sentence is very long, or the words relation inside the sentence are not tight, for example, there are some separators in the sentence, the model can get wrong translation. Aiming to solve these problems, in this paper, we propose a history attention structure that takes advantage of translated information. This architecture easily captures history information, helps model alleviate the memory vanishing problem introduced by long sentences and avoid focusing on one local part. In experiments, we show our history attention with gate improves translation quality.
AB - Attention mechanism has enhanced state-of-the-art Neural Machine Translation (NMT) by focusing on parts of the source sentence when predicting each target word. However we find that most of the attention context vector calculation is directly dependent on the current decoder hidden state. It tends to ignore past translated information, which often leads to over-translation and under-translation. When target sentence is very long, or the words relation inside the sentence are not tight, for example, there are some separators in the sentence, the model can get wrong translation. Aiming to solve these problems, in this paper, we propose a history attention structure that takes advantage of translated information. This architecture easily captures history information, helps model alleviate the memory vanishing problem introduced by long sentences and avoid focusing on one local part. In experiments, we show our history attention with gate improves translation quality.
KW - Attention
KW - Gate
KW - History Attention
KW - NMT
UR - https://www.scopus.com/pages/publications/85049799429
U2 - 10.1109/ICACI.2018.8377531
DO - 10.1109/ICACI.2018.8377531
M3 - 会议稿件
AN - SCOPUS:85049799429
T3 - Proceedings - 2018 10th International Conference on Advanced Computational Intelligence, ICACI 2018
SP - 619
EP - 624
BT - Proceedings - 2018 10th International Conference on Advanced Computational Intelligence, ICACI 2018
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 10th International Conference on Advanced Computational Intelligence, ICACI 2018
Y2 - 29 March 2018 through 31 March 2018
ER -