Abstract
Vision-Language-Action models (VLA), adapted from large-scale Vision-Language Models, have shown considerable promise for general-purpose robotic manipulation. However, we argue that these models often exhibit action-level hallucinations—generating actions that are inconsistent with visual observations—due to limitations inherited from their VLM backbones. Existing hallucination mitigation methods, such as Contrastive Decoding, are primarily designed for deterministic language tasks and are not well-suited to robotics, where multiple valid action trajectories may exist. To address this challenge, we propose Entropy-Optimized Contrastive Decoding (EOCD), a general decoding framework tailored for VLA models in robotic control. EOCD introduces a contrastive decoding formulation that balances probability mass among feasible action paths, thereby mitigating overconfidence in a single trajectory. Additionally, it incorporates an entropy-based optimization strategy that adaptively tunes decoding hyperparameters according to the maximum entropy principle, enhancing robustness in long-horizon predictions. Empirical results on public robotic manipulation benchmarks using the mainstream OpenVLA model demonstrate that EOCD significantly improves task success rates and robustness without requiring extra training.
| Original language | English |
|---|---|
| Article number | 132507 |
| Journal | Neurocomputing |
| Volume | 669 |
| DOIs | |
| State | Published - 7 Mar 2026 |
Keywords
- Contrastive decoding
- Hallucination suppression
- Maximum entropy
- Robotic manipulation
- Vision-language-action
Fingerprint
Dive into the research topics of 'Entropy-optimized contrastive decoding for hallucination suppression in vision-language-action models'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver