TY - JOUR
T1 - Beyond One-Size-Fits-All
T2 - Inversion Learning for Highly Effective NLG Evaluation Prompts
AU - Hong, Hanhua
AU - Xiao, Chenghao
AU - Wang, Yang
AU - Liu, Yiqi
AU - Rong, Wenge
AU - Lin, Chenghua
N1 - Publisher Copyright:
© 2026 Association for Computational Linguistics. This is an open-access article distributed under the terms of the https://creativecommons.org/licenses/by/4.0/, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. For a full description of the license, please visit https://creativecommons.org/licenses/by/4.0/legalcode.
PY - 2026
Y1 - 2026
N2 - Evaluating natural language generation systems is challenging due to the diversity of valid outputs. While human evaluation is the gold standard, it suffers from inconsistencies, lack of standardization, and demographic biases, limiting reproducibility. LLM-based evaluators offer a scalable alternative but are highly sensitive to prompt design, where small variations can lead to significant discrepancies. In this work, we propose an inversion learning method that learns effective reverse mappings from model outputs back to their input instructions, enabling the automatic generation of highly effective, model-specific evaluation prompts. Our method requires only a single evaluation sample and eliminates the need for time-consuming manual prompt engineering, thereby improving both efficiency and robustness. Our work contributes toward a new direction for more robust and efficient LLM-based evaluation.
AB - Evaluating natural language generation systems is challenging due to the diversity of valid outputs. While human evaluation is the gold standard, it suffers from inconsistencies, lack of standardization, and demographic biases, limiting reproducibility. LLM-based evaluators offer a scalable alternative but are highly sensitive to prompt design, where small variations can lead to significant discrepancies. In this work, we propose an inversion learning method that learns effective reverse mappings from model outputs back to their input instructions, enabling the automatic generation of highly effective, model-specific evaluation prompts. Our method requires only a single evaluation sample and eliminates the need for time-consuming manual prompt engineering, thereby improving both efficiency and robustness. Our work contributes toward a new direction for more robust and efficient LLM-based evaluation.
UR - https://www.scopus.com/pages/publications/105042057568
U2 - 10.1162/TACL.a.617
DO - 10.1162/TACL.a.617
M3 - 文章
AN - SCOPUS:105042057568
SN - 2307-387X
JO - Transactions of the Association for Computational Linguistics
JF - Transactions of the Association for Computational Linguistics
ER -