跳到主要导航 跳到搜索 跳到主要内容

IndVisSGG: VLM-based scene graph generation for industrial spatial intelligence

  • Beihang University
  • Nanyang Technological University
  • Huazhong University of Science and Technology

科研成果: 期刊稿件文章同行评审

摘要

Industrial spatial intelligence enables robots and machine tools to understand environmental settings and their relationships, allowing them to manipulate target components. A crucial aspect of this process is scene graph generation (SGG). Previous research on SGG primarily focuses on detection and panoptic segmentation of objects, followed by the prediction of their pairwise relationships. However, these approaches struggle with generalization and transferability when encountering new scenarios. To tackle this problem, we propose the Industrial Visual Scene Graph Generation (IndVisSGG) method, which parses spatial and interactive relationships between objects in temporal industrial settings. This approach leverages the capabilities of Vision-Language Models (VLMs) to generate scene graphs quickly and accurately without any additional object annotations. Furthermore, leveraging the IndVisSGG method, we have implemented a meticulous annotation procedure to compile a high-quality industrial scene graph generation (ISG) dataset, comprising 10,000 images of manufacturing and related industrial scenes. Through comparisons with various scene graph generation methods and benchmarks across two other datasets, we have showcased the superiority of the IndVisSGG method and underscored the benefits of the ISG dataset over existing datasets.

源语言英语
文章编号103107
期刊Advanced Engineering Informatics
65
DOI
出版状态已出版 - 5月 2025

指纹

探究 'IndVisSGG: VLM-based scene graph generation for industrial spatial intelligence' 的科研主题。它们共同构成独一无二的指纹。

引用此