Skip to main navigation Skip to search Skip to main content

CoSE: connectivity-oriented semantic enhancement for mitigating hallucinations in multimodal LLMs

  • Yuanze Hu
  • , Zhaoxin Fan
  • , Gen Li
  • , Zhichao Yang
  • , Xinyu Wang
  • , Ye Qiu
  • , Wenjun Wu
  • , Kejian Wu
  • , Yifan Sun
  • , Xiaotie Deng
  • , Jin Dong*
  • , Ziyu Jia
  • *Corresponding author for this work
  • Beihang University
  • Xreal
  • Renmin University of China
  • Peking University
  • Beijing Academy of Blockchain and Edge Computing
  • Chinese Academy of Sciences

Research output: Contribution to journalArticlepeer-review

Abstract

Multimodal Large Language Models (MLLMs) have achieved remarkable progress, yet their lightweight variants remain highly susceptible to hallucinations-generating outputs inconsistent with visual inputs. While empirical mitigation strategies have been proposed, the fundamental question of why smaller models hallucinate more remains poorly understood. Therefore, this paper provides the first systematic investigation through a dynamical systems lens, suggesting that hallucination propensity is intrinsically linked to the geometric structure of semantic manifolds. Through comprehensive multi-scale, multi-difficulty analysis, we uncover three critical findings: (i) Smaller models exhibit weaker manifold connectivity; (ii) As task difficulty increases, manifold connectivity weakens; (iii) Weaker manifold connectivity is associated with deeper attractor escapes. These findings suggest that weak semantic connectivity is an important geometric factor associated with hallucination susceptibility in lightweight MLLMs. Motivated by this empirical insight, we propose CoSE (Connectivity-Oriented Semantic Enhancement), a training-inference consistent framework that augments representations with retrieved semantically connected latent samples. While keeping the main architecture unchanged and incurring negligible overhead, CoSE consistently reduces hallucinations and boosts performance across diverse VQA benchmarks via a small plug-in module.

Original languageEnglish
Article number104478
JournalInformation Fusion
Volume135
DOIs
StatePublished - Nov 2026

Keywords

  • Dynamical systems
  • Hallucination mitigation
  • Multimodal large language models
  • Semantic manifold

Fingerprint

Dive into the research topics of 'CoSE: connectivity-oriented semantic enhancement for mitigating hallucinations in multimodal LLMs'. Together they form a unique fingerprint.

Cite this