跳到主要导航 跳到搜索 跳到主要内容

PromDA: Prompt-based Data Augmentation for Low-Resource NLU Tasks

  • Yufei Wang
  • , Can Xu
  • , Qingfeng Sun
  • , Huang Hu
  • , Chongyang Tao
  • , Xiubo Geng
  • , Daxin Jiang
  • Macquarie University
  • Microsoft USA
  • Microsoft

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

This paper focuses on the Data Augmentation for low-resource Natural Language Understanding (NLU) tasks. We propose Prompt-based Data Augmentation model (PromDA) which only trains small-scale Soft Prompt (i.e., a set of trainable vectors) in the frozen Pre-trained Language Models (PLMs). This avoids human effort in collecting unlabeled in-domain data and maintains the quality of generated synthetic data. In addition, PromDA generates synthetic data via two different views and filters out the low-quality data using NLU models. Experiments on four benchmarks show that synthetic data produced by PromDA successfully boost up the performance of NLU models which consistently outperform several competitive baseline models, including a state-of-the-art semi-supervised model using unlabeled in-domain data. The synthetic data from PromDA are also complementary with unlabeled in-domain data. The NLU models can be further improved when they are combined for training.

源语言英语
主期刊名ACL 2022 - 60th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference (Long Papers)
编辑Smaranda Muresan, Preslav Nakov, Aline Villavicencio
出版商Association for Computational Linguistics (ACL)
4242-4255
页数14
ISBN(电子版)9781955917216
DOI
出版状态已出版 - 2022
已对外发布
活动60th Annual Meeting of the Association for Computational Linguistics, ACL 2022 - Dublin, 爱尔兰
期限: 22 5月 202227 5月 2022

出版系列

姓名Proceedings of the Annual Meeting of the Association for Computational Linguistics
1
ISSN(印刷版)0736-587X

会议

会议60th Annual Meeting of the Association for Computational Linguistics, ACL 2022
国家/地区爱尔兰
Dublin
时期22/05/2227/05/22

学术指纹

探究 'PromDA: Prompt-based Data Augmentation for Low-Resource NLU Tasks' 的科研主题。它们共同构成独一无二的学术指纹。

引用此