摘要
Multimodal Large Language Models (MLLMs) have achieved remarkable progress, yet their lightweight variants remain highly susceptible to hallucinations-generating outputs inconsistent with visual inputs. While empirical mitigation strategies have been proposed, the fundamental question of why smaller models hallucinate more remains poorly understood. Therefore, this paper provides the first systematic investigation through a dynamical systems lens, suggesting that hallucination propensity is intrinsically linked to the geometric structure of semantic manifolds. Through comprehensive multi-scale, multi-difficulty analysis, we uncover three critical findings: (i) Smaller models exhibit weaker manifold connectivity; (ii) As task difficulty increases, manifold connectivity weakens; (iii) Weaker manifold connectivity is associated with deeper attractor escapes. These findings suggest that weak semantic connectivity is an important geometric factor associated with hallucination susceptibility in lightweight MLLMs. Motivated by this empirical insight, we propose CoSE (Connectivity-Oriented Semantic Enhancement), a training-inference consistent framework that augments representations with retrieved semantically connected latent samples. While keeping the main architecture unchanged and incurring negligible overhead, CoSE consistently reduces hallucinations and boosts performance across diverse VQA benchmarks via a small plug-in module.
| 源语言 | 英语 |
|---|---|
| 期刊论文编号 | 104478 |
| 期刊 | Information Fusion |
| 卷 | 135 |
| DOI | |
| 出版状态 | 已出版 - 11月 2026 |
学术指纹
探究 'CoSE: connectivity-oriented semantic enhancement for mitigating hallucinations in multimodal LLMs' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver