Abstract
Industrial spatial intelligence enables robots and machine tools to understand environmental settings and their relationships, allowing them to manipulate target components. A crucial aspect of this process is scene graph generation (SGG). Previous research on SGG primarily focuses on detection and panoptic segmentation of objects, followed by the prediction of their pairwise relationships. However, these approaches struggle with generalization and transferability when encountering new scenarios. To tackle this problem, we propose the Industrial Visual Scene Graph Generation (IndVisSGG) method, which parses spatial and interactive relationships between objects in temporal industrial settings. This approach leverages the capabilities of Vision-Language Models (VLMs) to generate scene graphs quickly and accurately without any additional object annotations. Furthermore, leveraging the IndVisSGG method, we have implemented a meticulous annotation procedure to compile a high-quality industrial scene graph generation (ISG) dataset, comprising 10,000 images of manufacturing and related industrial scenes. Through comparisons with various scene graph generation methods and benchmarks across two other datasets, we have showcased the superiority of the IndVisSGG method and underscored the benefits of the ISG dataset over existing datasets.
| Original language | English |
|---|---|
| Article number | 103107 |
| Journal | Advanced Engineering Informatics |
| Volume | 65 |
| DOIs | |
| State | Published - May 2025 |
Keywords
- Industrial spatial intelligence
- Scene graph
- Smart manufacturing
- Vision-language model
Fingerprint
Dive into the research topics of 'IndVisSGG: VLM-based scene graph generation for industrial spatial intelligence'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver