TY - JOUR
T1 - GenderBias-VL
T2 - Benchmarking Gender Bias in Vision Language Models via Counterfactual Probing
AU - Xiao, Yisong
AU - Liu, Xianglong
AU - Cheng, Qian Jia
AU - Yin, Zhenfei
AU - Liang, Siyuan
AU - Li, Jiapeng
AU - Shao, Jing
AU - Liu, Aishan
AU - Tao, Dacheng
N1 - Publisher Copyright:
© The Author(s), under exclusive licence to Springer Science+Business Media, LLC, part of Springer Nature 2025.
PY - 2025/12
Y1 - 2025/12
N2 - Large Vision-Language Models (LVLMs) have been widely adopted in various applications; however, they exhibit significant gender biases. Existing benchmarks primarily evaluate gender bias at the demographic group level, neglecting individual fairness, which emphasizes equal treatment of similar individuals. This research gap limits the detection of discriminatory behaviors, as individual fairness offers a more granular examination of biases that group fairness may overlook. For the first time, this paper introduces the GenderBias-VL benchmark to evaluate occupation-related gender bias in LVLMs using counterfactual visual questions under individual fairness criteria. To construct this benchmark, we first utilize text-to-image diffusion models to generate occupation images and their gender counterfactuals. Subsequently, we generate corresponding textual occupation options by identifying stereotyped occupation pairs with high semantic similarity but opposite gender proportions in real-world statistics. This method enables the creation of large-scale visual question counterfactuals to expose biases in LVLMs, applicable in both multimodal and unimodal contexts through modifying gender attributes in specific modalities. Overall, our GenderBias-VL benchmark comprises 34,581 visual question counterfactual pairs, covering 177 occupations. Using our benchmark, we extensively evaluate 19 commonly used open-source LVLMs (e.g., LLaVA) and state-of-the-art commercial APIs (e.g., GPT and Gemini). Our findings reveal widespread gender biases in existing LVLMs. Our benchmark offers: (1) a comprehensive dataset for occupation-related gender bias evaluation; (2) an up-to-date leaderboard on LVLM biases; and (3) a nuanced understanding of the biases presented by these models. The dataset and code are available at the https://genderbiasvl.github.io.
AB - Large Vision-Language Models (LVLMs) have been widely adopted in various applications; however, they exhibit significant gender biases. Existing benchmarks primarily evaluate gender bias at the demographic group level, neglecting individual fairness, which emphasizes equal treatment of similar individuals. This research gap limits the detection of discriminatory behaviors, as individual fairness offers a more granular examination of biases that group fairness may overlook. For the first time, this paper introduces the GenderBias-VL benchmark to evaluate occupation-related gender bias in LVLMs using counterfactual visual questions under individual fairness criteria. To construct this benchmark, we first utilize text-to-image diffusion models to generate occupation images and their gender counterfactuals. Subsequently, we generate corresponding textual occupation options by identifying stereotyped occupation pairs with high semantic similarity but opposite gender proportions in real-world statistics. This method enables the creation of large-scale visual question counterfactuals to expose biases in LVLMs, applicable in both multimodal and unimodal contexts through modifying gender attributes in specific modalities. Overall, our GenderBias-VL benchmark comprises 34,581 visual question counterfactual pairs, covering 177 occupations. Using our benchmark, we extensively evaluate 19 commonly used open-source LVLMs (e.g., LLaVA) and state-of-the-art commercial APIs (e.g., GPT and Gemini). Our findings reveal widespread gender biases in existing LVLMs. Our benchmark offers: (1) a comprehensive dataset for occupation-related gender bias evaluation; (2) an up-to-date leaderboard on LVLM biases; and (3) a nuanced understanding of the biases presented by these models. The dataset and code are available at the https://genderbiasvl.github.io.
KW - Benchmark
KW - Gender Bias
KW - Individual Fairness
KW - Large Vision Language Models
UR - https://www.scopus.com/pages/publications/105016814884
U2 - 10.1007/s11263-025-02556-7
DO - 10.1007/s11263-025-02556-7
M3 - 文章
AN - SCOPUS:105016814884
SN - 0920-5691
VL - 133
SP - 8332
EP - 8355
JO - International Journal of Computer Vision
JF - International Journal of Computer Vision
IS - 12
ER -