Skip to main navigation Skip to search Skip to main content

Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts

  • Hanhua Hong
  • , Chenghao Xiao
  • , Yang Wang
  • , Yiqi Liu
  • , Wenge Rong
  • , Chenghua Lin*
  • *Corresponding author for this work
  • University of Manchester
  • Durham University

Research output: Contribution to journalArticlepeer-review

Abstract

Evaluating natural language generation systems is challenging due to the diversity of valid outputs. While human evaluation is the gold standard, it suffers from inconsistencies, lack of standardization, and demographic biases, limiting reproducibility. LLM-based evaluators offer a scalable alternative but are highly sensitive to prompt design, where small variations can lead to significant discrepancies. In this work, we propose an inversion learning method that learns effective reverse mappings from model outputs back to their input instructions, enabling the automatic generation of highly effective, model-specific evaluation prompts. Our method requires only a single evaluation sample and eliminates the need for time-consuming manual prompt engineering, thereby improving both efficiency and robustness. Our work contributes toward a new direction for more robust and efficient LLM-based evaluation.

Original languageEnglish
Pages (from-to)689-710
Number of pages22
JournalTransactions of the Association for Computational Linguistics
Volume14
DOIs
StatePublished - 2026

Fingerprint

Dive into the research topics of 'Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts'. Together they form a unique fingerprint.

Cite this