TY - JOUR
T1 - MindScore
T2 - quantifying human preference for text-to-image generation through multi-view lens
AU - Tong, Yiqi
AU - Zhang, Jiarui
AU - Wei, Shaohang
AU - Guo, Wei
AU - Zhuang, Fuzhen
AU - Wang, Deqing
AU - Yang, Xi
AU - Xuan, Richeng
N1 - Publisher Copyright:
© Science China Press 2025.
PY - 2025/6
Y1 - 2025/6
N2 - Understanding and quantifying the capabilities of foundation models, particularly in text-to-image (T2I) generation, is crucial for verifying their alignment with human expectations and practical requirements. However, evaluating T2I foundation models presents significant challenges due to the complex, multi-dimensional psychological factors that influence human preferences for generated images. In this work, we propose MindScore, a multi-view framework for assessing the generation capacity of T2I models through the lens of human preference. Specifically, MindScore decomposes the evaluation into four complementary modules that align with human cognitive processing of images: matching, faithfulness, quality, and realness. The matching module quantifies the semantic alignment between generated images and prompt text, while the faithfulness module measures how accurately the images reflect specific prompt details. Furthermore, we incorporate quality and realness modules to capture deeper psychological preferences, recognizing that unpleasant or distorted images often trigger adverse human responses. Extensive experiments on three T2I datasets with human preference annotations clearly validate the superiority of our proposed MindScore over various state-of-the-art baselines. Our case studies further reveal that MindScore offers valuable insights into T2I generation from a human-centric perspective.
AB - Understanding and quantifying the capabilities of foundation models, particularly in text-to-image (T2I) generation, is crucial for verifying their alignment with human expectations and practical requirements. However, evaluating T2I foundation models presents significant challenges due to the complex, multi-dimensional psychological factors that influence human preferences for generated images. In this work, we propose MindScore, a multi-view framework for assessing the generation capacity of T2I models through the lens of human preference. Specifically, MindScore decomposes the evaluation into four complementary modules that align with human cognitive processing of images: matching, faithfulness, quality, and realness. The matching module quantifies the semantic alignment between generated images and prompt text, while the faithfulness module measures how accurately the images reflect specific prompt details. Furthermore, we incorporate quality and realness modules to capture deeper psychological preferences, recognizing that unpleasant or distorted images often trigger adverse human responses. Extensive experiments on three T2I datasets with human preference annotations clearly validate the superiority of our proposed MindScore over various state-of-the-art baselines. Our case studies further reveal that MindScore offers valuable insights into T2I generation from a human-centric perspective.
KW - foundation models
KW - human preference evaluation
KW - language and vision
KW - multi-view assessment
KW - text-to-image generation
UR - https://www.scopus.com/pages/publications/105006915316
U2 - 10.1007/s11432-024-4401-y
DO - 10.1007/s11432-024-4401-y
M3 - 文章
AN - SCOPUS:105006915316
SN - 1674-733X
VL - 68
JO - Science China Information Sciences
JF - Science China Information Sciences
IS - 6
M1 - 160105
ER -