TY - JOUR
T1 - Entropy-optimized contrastive decoding for hallucination suppression in vision-language-action models
AU - Qiu, Ye
AU - Fan, Zhaoxin
AU - Yu, Qingchen
AU - Wu, Faguo
AU - Zheng, Hongwei
AU - Sun, Yifan
AU - Wu, Wenjun
N1 - Publisher Copyright:
© 2025 Elsevier B.V.
PY - 2026/3/7
Y1 - 2026/3/7
N2 - Vision-Language-Action models (VLA), adapted from large-scale Vision-Language Models, have shown considerable promise for general-purpose robotic manipulation. However, we argue that these models often exhibit action-level hallucinations—generating actions that are inconsistent with visual observations—due to limitations inherited from their VLM backbones. Existing hallucination mitigation methods, such as Contrastive Decoding, are primarily designed for deterministic language tasks and are not well-suited to robotics, where multiple valid action trajectories may exist. To address this challenge, we propose Entropy-Optimized Contrastive Decoding (EOCD), a general decoding framework tailored for VLA models in robotic control. EOCD introduces a contrastive decoding formulation that balances probability mass among feasible action paths, thereby mitigating overconfidence in a single trajectory. Additionally, it incorporates an entropy-based optimization strategy that adaptively tunes decoding hyperparameters according to the maximum entropy principle, enhancing robustness in long-horizon predictions. Empirical results on public robotic manipulation benchmarks using the mainstream OpenVLA model demonstrate that EOCD significantly improves task success rates and robustness without requiring extra training.
AB - Vision-Language-Action models (VLA), adapted from large-scale Vision-Language Models, have shown considerable promise for general-purpose robotic manipulation. However, we argue that these models often exhibit action-level hallucinations—generating actions that are inconsistent with visual observations—due to limitations inherited from their VLM backbones. Existing hallucination mitigation methods, such as Contrastive Decoding, are primarily designed for deterministic language tasks and are not well-suited to robotics, where multiple valid action trajectories may exist. To address this challenge, we propose Entropy-Optimized Contrastive Decoding (EOCD), a general decoding framework tailored for VLA models in robotic control. EOCD introduces a contrastive decoding formulation that balances probability mass among feasible action paths, thereby mitigating overconfidence in a single trajectory. Additionally, it incorporates an entropy-based optimization strategy that adaptively tunes decoding hyperparameters according to the maximum entropy principle, enhancing robustness in long-horizon predictions. Empirical results on public robotic manipulation benchmarks using the mainstream OpenVLA model demonstrate that EOCD significantly improves task success rates and robustness without requiring extra training.
KW - Contrastive decoding
KW - Hallucination suppression
KW - Maximum entropy
KW - Robotic manipulation
KW - Vision-language-action
UR - https://www.scopus.com/pages/publications/105026144090
U2 - 10.1016/j.neucom.2025.132507
DO - 10.1016/j.neucom.2025.132507
M3 - 文章
AN - SCOPUS:105026144090
SN - 0925-2312
VL - 669
JO - Neurocomputing
JF - Neurocomputing
M1 - 132507
ER -