Abstract
Multimodal Large Language Models (MLLMs) have achieved remarkable progress, yet their lightweight variants remain highly susceptible to hallucinations-generating outputs inconsistent with visual inputs. While empirical mitigation strategies have been proposed, the fundamental question of why smaller models hallucinate more remains poorly understood. Therefore, this paper provides the first systematic investigation through a dynamical systems lens, suggesting that hallucination propensity is intrinsically linked to the geometric structure of semantic manifolds. Through comprehensive multi-scale, multi-difficulty analysis, we uncover three critical findings: (i) Smaller models exhibit weaker manifold connectivity; (ii) As task difficulty increases, manifold connectivity weakens; (iii) Weaker manifold connectivity is associated with deeper attractor escapes. These findings suggest that weak semantic connectivity is an important geometric factor associated with hallucination susceptibility in lightweight MLLMs. Motivated by this empirical insight, we propose CoSE (Connectivity-Oriented Semantic Enhancement), a training-inference consistent framework that augments representations with retrieved semantically connected latent samples. While keeping the main architecture unchanged and incurring negligible overhead, CoSE consistently reduces hallucinations and boosts performance across diverse VQA benchmarks via a small plug-in module.
| Original language | English |
|---|---|
| Article number | 104478 |
| Journal | Information Fusion |
| Volume | 135 |
| DOIs | |
| State | Published - Nov 2026 |
Keywords
- Dynamical systems
- Hallucination mitigation
- Multimodal large language models
- Semantic manifold
Fingerprint
Dive into the research topics of 'CoSE: connectivity-oriented semantic enhancement for mitigating hallucinations in multimodal LLMs'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver