跳到主要导航 跳到搜索 跳到主要内容

Improving the Robustness of Summarization Systems with Dual Augmentation

  • Xiuying Chen
  • , Guodong Long
  • , Chongyang Tao*
  • , Mingzhe Li
  • , Xin Gao*
  • , Chengqi Zhang
  • , Xiangliang Zhang
  • *此作品的通讯作者
  • King Abdullah University of Science and Technology
  • University of Technology Sydney
  • Microsoft USA
  • Ant Group
  • University of Notre Dame

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

A robust summarization system should be able to capture the gist of the document, regardless of the specific word choices or noise in the input. In this work, we first explore the summarization models' robustness against perturbations including word-level synonym substitution and noise. To create semantic-consistent substitutes, we propose a SummAttacker, which is an efficient approach to generating adversarial samples based on language models. Experimental results show that state-of-the-art summarization models have a significant decrease in performance on adversarial and noisy test sets. Next, we analyze the vulnerability of the summarization systems and explore improving the robustness by data augmentation. Specifically, the first brittleness factor we found is the poor understanding of infrequent words in the input. Correspondingly, we feed the encoder with more diverse cases created by SummAttacker in the input space. The other factor is in the latent space, where the attacked inputs bring more variations to the hidden states. Hence, we construct adversarial decoder input and devise manifold softmixing operation in hidden space to introduce more diversity. Experimental results on Gigaword and CNN/DM datasets demonstrate that our approach achieves significant improvements over strong baselines and exhibits higher robustness on noisy, attacked, and clean datasets.

源语言英语
主期刊名Long Papers
出版商Association for Computational Linguistics (ACL)
6846-6857
页数12
ISBN(电子版)9781959429722
DOI
出版状态已出版 - 2023
已对外发布
活动61st Annual Meeting of the Association for Computational Linguistics, ACL 2023 - Toronto, 加拿大
期限: 9 7月 202314 7月 2023

出版系列

姓名Proceedings of the Annual Meeting of the Association for Computational Linguistics
1
ISSN(印刷版)0736-587X

会议

会议61st Annual Meeting of the Association for Computational Linguistics, ACL 2023
国家/地区加拿大
Toronto
时期9/07/2314/07/23

指纹

探究 'Improving the Robustness of Summarization Systems with Dual Augmentation' 的科研主题。它们共同构成独一无二的指纹。

引用此