跳到主要导航 跳到搜索 跳到主要内容

Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment

  • Hao Li
  • , Lijun Li*
  • , Zhenghao Lu
  • , Xianyi Wei
  • , Rui Li
  • , Jing Shao*
  • , Lei Sha*
  • *此作品的通讯作者
  • Shanghai Artificial Intelligence Laboratory
  • Beihang University
  • Wuhan University
  • Peking University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

With rapid advancement and increasing accessibility of LLMs, fine-tuning aligned models has become a critical step for adapting them to real-world applications, which makes the safety of this fine-tuning process more important than ever. However, recent studies have highlighted a critical challenge: even when fine-tuning with benign datasets, the safety alignment of aligned LLMs can be compromised, making them more susceptible to malicious instructions. In this paper, we show that fine-tuning datasets often contain safety-degrading samples that are not easily identifiable on the surface. These samples can easily degrade the safety alignment of LLMs during fine-tuning. To address this issue, we propose LARF, a Layer-Aware Representation Filtering method. This method identifies safety-sensitive layers within the LLM and leverages data representations to detect safety-degrading data samples in the fine-tuning dataset. Experimental results demonstrate that LARF can efficiently and effectively identify safety-degrading data. After removing such data, the safety alignment degradation caused by fine-tuning is mitigated. Please see our code at https://github.com/LLLeoLi/LARF.

源语言英语
主期刊名EMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference
编辑Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng
出版商Association for Computational Linguistics (ACL)
8030-8050
页数21
ISBN(电子版)9798891763326
DOI
出版状态已出版 - 2025
活动30th Conference on Empirical Methods in Natural Language Processing, EMNLP 2025 - Suzhou, 中国
期限: 4 11月 20259 11月 2025

出版系列

姓名EMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference

会议

会议30th Conference on Empirical Methods in Natural Language Processing, EMNLP 2025
国家/地区中国
Suzhou
时期4/11/259/11/25

指纹

探究 'Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment' 的科研主题。它们共同构成独一无二的指纹。

引用此