跳到主要导航 跳到搜索 跳到主要内容

Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts

  • Hanhua Hong
  • , Chenghao Xiao
  • , Yang Wang
  • , Yiqi Liu
  • , Wenge Rong
  • , Chenghua Lin*
  • *此作品的通讯作者
  • University of Manchester
  • Durham University

科研成果: 期刊稿件文章同行评审

摘要

Evaluating natural language generation systems is challenging due to the diversity of valid outputs. While human evaluation is the gold standard, it suffers from inconsistencies, lack of standardization, and demographic biases, limiting reproducibility. LLM-based evaluators offer a scalable alternative but are highly sensitive to prompt design, where small variations can lead to significant discrepancies. In this work, we propose an inversion learning method that learns effective reverse mappings from model outputs back to their input instructions, enabling the automatic generation of highly effective, model-specific evaluation prompts. Our method requires only a single evaluation sample and eliminates the need for time-consuming manual prompt engineering, thereby improving both efficiency and robustness. Our work contributes toward a new direction for more robust and efficient LLM-based evaluation.

学术指纹

探究 'Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts' 的科研主题。它们共同构成独一无二的学术指纹。

引用此