跳到主要导航 跳到搜索 跳到主要内容

Enhancing open-vocabulary scene understanding via push–pull alignment in gaussian splatting

  • Tong Chen
  • , Shengjia Liang
  • , Yuan Xiong
  • , Qiang Zhou
  • , Qichuan Geng
  • , Zhong Zhou*
  • *此作品的通讯作者
  • Beihang University
  • Sun Yat-Sen University
  • Capital Normal University
  • Zhongguancun Laboratory

科研成果: 期刊稿件文章同行评审

摘要

Open-vocabulary scene understanding based on 3D Gaussian Splatting (3DGS) has shown promising potential for applications such as embodied agents and object localization. By integrating open-vocabulary embeddings into spatial 3D gaussians, these models enable a more comprehensive understanding of scenes. However, existing methods often suffer from misalignment due to the gap between RGB and language modalities, leading to incorrect interpretations of similar-looking objects. To address this issue, we propose a cross-modal integration approach that aligns multiple representations through spatial gaussian positioning. We introduce Push-Pull alignment in Gaussian Splatting(PPGS), a novel bimodal framework that bridges RGB and language modalities through cohesive representation fields. Leveraging the illumination-invariant properties of language embeddings, we design the bridge module, which uses the geometrically-grounded positions for the gaussians as a direct bridge between the two modalities. This module significantly enhances cross-modal alignment, improves high-fidelity rendering, and ensures accurate language feature embeddings. Furthermore, our framework dynamically adjusts gradients based on the distinct optimization requirements of RGB and language during joint learning, ensuring stable and efficient convergence. Comprehensive experiments demonstrate that PPGS achieves superior language query accuracy and enhanced visual quality compared to existing language-embedded representations, with Intersection over Union (mIoU) increasing by 6% and Peak Signal-to-Noise Ratio (PSNR) showing gains over mainstream methods, all within only 50% of the training time.

源语言英语
文章编号38
期刊Visual Computer
42
1
DOI
出版状态已出版 - 1月 2026

学术指纹

探究 'Enhancing open-vocabulary scene understanding via push–pull alignment in gaussian splatting' 的科研主题。它们共同构成独一无二的学术指纹。

引用此