Skip to main navigation Skip to search Skip to main content

MindScore: quantifying human preference for text-to-image generation through multi-view lens

  • Yiqi Tong
  • , Jiarui Zhang
  • , Shaohang Wei
  • , Wei Guo
  • , Fuzhen Zhuang*
  • , Deqing Wang
  • , Xi Yang
  • , Richeng Xuan
  • *Corresponding author for this work
  • Beihang University
  • Shanghai Jiao Tong University
  • Peking University
  • Beijing Academy of Artificial Intelligence

Research output: Contribution to journalArticlepeer-review

Abstract

Understanding and quantifying the capabilities of foundation models, particularly in text-to-image (T2I) generation, is crucial for verifying their alignment with human expectations and practical requirements. However, evaluating T2I foundation models presents significant challenges due to the complex, multi-dimensional psychological factors that influence human preferences for generated images. In this work, we propose MindScore, a multi-view framework for assessing the generation capacity of T2I models through the lens of human preference. Specifically, MindScore decomposes the evaluation into four complementary modules that align with human cognitive processing of images: matching, faithfulness, quality, and realness. The matching module quantifies the semantic alignment between generated images and prompt text, while the faithfulness module measures how accurately the images reflect specific prompt details. Furthermore, we incorporate quality and realness modules to capture deeper psychological preferences, recognizing that unpleasant or distorted images often trigger adverse human responses. Extensive experiments on three T2I datasets with human preference annotations clearly validate the superiority of our proposed MindScore over various state-of-the-art baselines. Our case studies further reveal that MindScore offers valuable insights into T2I generation from a human-centric perspective.

Original languageEnglish
Article number160105
JournalScience China Information Sciences
Volume68
Issue number6
DOIs
StatePublished - Jun 2025

Keywords

  • foundation models
  • human preference evaluation
  • language and vision
  • multi-view assessment
  • text-to-image generation

Fingerprint

Dive into the research topics of 'MindScore: quantifying human preference for text-to-image generation through multi-view lens'. Together they form a unique fingerprint.

Cite this