Abstract
Conventional fine-tuning encounters increasing difficulties given the size of current Pre-trained Language Models, which makes parameter-efficient tuning become the focal point of frontier research. Recent advances in this field is the unified tuning methods that aim to tune the representations of both multi-head attention (MHA) and fully connected feed-forward network (FFN) simultaneously, but they rely on existing tuning methods and do not explicitly model domain knowledge for downstream tasks. In this work, we propose memory-tuning, a novel unified parameter-efficient tuning method with task-specific knowledge learning, for both MHA and FFN components in Transformer blocks. We also prove that the well-known prefix tuning is also a kind of memory tuning, which further ensures memory tuning is a genuine unified tuning method. Experiments on eight benchmark data sets including both sentence- and token-level tasks demonstrate that our method outperforms the state-of-the-art baselines even full-tuning in most cases.
| Original language | English |
|---|---|
| Journal | IEEE/ACM Transactions on Audio Speech and Language Processing |
| DOIs | |
| State | Accepted/In press - 2024 |
Keywords
- Memory Mechanism
- Parameter-Efficient Tuning
- Pre-trained Language Model
Fingerprint
Dive into the research topics of 'Memory-Tuning: A Unified Parameter-Efficient Tuning Method for Pre-trained Language Models'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver