TY - GEN
T1 - GANLM
T2 - 61st Annual Meeting of the Association for Computational Linguistics, ACL 2023
AU - Yang, Jian
AU - Ma, Shuming
AU - Dong, Li
AU - Huang, Shaohan
AU - Huang, Haoyang
AU - Yin, Yuwei
AU - Zhang, Dongdong
AU - Yang, Liqun
AU - Wei, Furu
AU - Li, Zhoujun
N1 - Publisher Copyright:
© 2023 Association for Computational Linguistics.
PY - 2023
Y1 - 2023
N2 - Pre-trained models have achieved remarkable success in natural language processing (NLP). However, existing pre-training methods under-utilize the benefits of language understanding for generation. Inspired by the idea of Generative Adversarial Networks (GANs), we propose a GAN-style model for encoder-decoder pretraining by introducing an auxiliary discriminator, unifying the ability of language understanding and generation in a single model. Our model, named as GANLM, is trained with two pre-training objectives: replaced token detection and replaced token denoising. Specifically, given masked source sentences, the generator outputs the target distribution and the discriminator predicts whether the target sampled tokens from distribution are incorrect. The target sentence is replaced with misclassified tokens to construct noisy previous context, which is used to generate the gold sentence. In general, both tasks improve the ability of language understanding and generation by selectively using the denoising data. Extensive experiments in language generation benchmarks show that GANLM with the powerful language understanding capability outperforms various strong pre-trained language models (PLMs) and achieves state-of-the-art performance.
AB - Pre-trained models have achieved remarkable success in natural language processing (NLP). However, existing pre-training methods under-utilize the benefits of language understanding for generation. Inspired by the idea of Generative Adversarial Networks (GANs), we propose a GAN-style model for encoder-decoder pretraining by introducing an auxiliary discriminator, unifying the ability of language understanding and generation in a single model. Our model, named as GANLM, is trained with two pre-training objectives: replaced token detection and replaced token denoising. Specifically, given masked source sentences, the generator outputs the target distribution and the discriminator predicts whether the target sampled tokens from distribution are incorrect. The target sentence is replaced with misclassified tokens to construct noisy previous context, which is used to generate the gold sentence. In general, both tasks improve the ability of language understanding and generation by selectively using the denoising data. Extensive experiments in language generation benchmarks show that GANLM with the powerful language understanding capability outperforms various strong pre-trained language models (PLMs) and achieves state-of-the-art performance.
UR - https://www.scopus.com/pages/publications/85171677161
U2 - 10.18653/v1/2023.acl-long.522
DO - 10.18653/v1/2023.acl-long.522
M3 - 会议稿件
AN - SCOPUS:85171677161
T3 - Proceedings of the Annual Meeting of the Association for Computational Linguistics
SP - 9394
EP - 9412
BT - Long Papers
PB - Association for Computational Linguistics (ACL)
Y2 - 9 July 2023 through 14 July 2023
ER -