跳到主要导航 跳到搜索 跳到主要内容

基于 VL 模型蒸馏与 LLM 解析的 三维场景图生成方法

  • Beihang University

科研成果: 期刊稿件文章同行评审

摘要

To address the limitation of point clouds in expressing semantic relationships for 3D scene-graph generation tasks, which typically requires rendering corresponding images and fusing multimodal features-thereby introducing additional computational overhead during inference, a 3D scene-graph generation method based on Vision-Language Model (VL model) distillation and Large Language Models (LLM) was proposed. The method took 3D point clouds as input, rendered corresponding images, and aligned their feature spaces to distill knowledge from the VL model into a Graph Neural Network (GNN), thereby establishing a mapping between point-cloud instances and corresponding textual descriptions and constructing a Point-cloud-Language model (PL model). The PL model leveraged an LLM to enhance the understanding of complex semantic relationships and effectively aggregated node features through the GNN. It could capture both semantic and spatial relationships of point clouds without relying on additional image information, enabling 3D scene-graph generation for indoor environments. Experimental results demonstrated that the proposed method not only achieved robust understanding of 3D indoor environments in open-vocabulary tasks, but also significantly reduced computational overhead and inference time compared with end-to-end 3D scene-graph generation approaches that relied on vision-language models, highlighting its strong performance and practical applicability.

投稿的翻译标题3D scene-graph generation via vision-language model distillation and large language model parsing
源语言繁体中文
页(从-至)360-367
页数8
期刊Journal of Graphics
47
2
DOI
出版状态已出版 - 30 4月 2026

关键词

  • 3D scene understanding
  • 3D scene-graph generation
  • knowledge distillation
  • large language model
  • vision-language model

学术指纹

探究 '基于 VL 模型蒸馏与 LLM 解析的 三维场景图生成方法' 的科研主题。它们共同构成独一无二的学术指纹。

引用此