TY - JOUR
T1 - IndVisSGG
T2 - VLM-based scene graph generation for industrial spatial intelligence
AU - Wang, Zuoxu
AU - Yan, Zhijie
AU - Li, Shufei
AU - Liu, Jihong
N1 - Publisher Copyright:
© 2025 Elsevier Ltd
PY - 2025/5
Y1 - 2025/5
N2 - Industrial spatial intelligence enables robots and machine tools to understand environmental settings and their relationships, allowing them to manipulate target components. A crucial aspect of this process is scene graph generation (SGG). Previous research on SGG primarily focuses on detection and panoptic segmentation of objects, followed by the prediction of their pairwise relationships. However, these approaches struggle with generalization and transferability when encountering new scenarios. To tackle this problem, we propose the Industrial Visual Scene Graph Generation (IndVisSGG) method, which parses spatial and interactive relationships between objects in temporal industrial settings. This approach leverages the capabilities of Vision-Language Models (VLMs) to generate scene graphs quickly and accurately without any additional object annotations. Furthermore, leveraging the IndVisSGG method, we have implemented a meticulous annotation procedure to compile a high-quality industrial scene graph generation (ISG) dataset, comprising 10,000 images of manufacturing and related industrial scenes. Through comparisons with various scene graph generation methods and benchmarks across two other datasets, we have showcased the superiority of the IndVisSGG method and underscored the benefits of the ISG dataset over existing datasets.
AB - Industrial spatial intelligence enables robots and machine tools to understand environmental settings and their relationships, allowing them to manipulate target components. A crucial aspect of this process is scene graph generation (SGG). Previous research on SGG primarily focuses on detection and panoptic segmentation of objects, followed by the prediction of their pairwise relationships. However, these approaches struggle with generalization and transferability when encountering new scenarios. To tackle this problem, we propose the Industrial Visual Scene Graph Generation (IndVisSGG) method, which parses spatial and interactive relationships between objects in temporal industrial settings. This approach leverages the capabilities of Vision-Language Models (VLMs) to generate scene graphs quickly and accurately without any additional object annotations. Furthermore, leveraging the IndVisSGG method, we have implemented a meticulous annotation procedure to compile a high-quality industrial scene graph generation (ISG) dataset, comprising 10,000 images of manufacturing and related industrial scenes. Through comparisons with various scene graph generation methods and benchmarks across two other datasets, we have showcased the superiority of the IndVisSGG method and underscored the benefits of the ISG dataset over existing datasets.
KW - Industrial spatial intelligence
KW - Scene graph
KW - Smart manufacturing
KW - Vision-language model
UR - https://www.scopus.com/pages/publications/85214567426
U2 - 10.1016/j.aei.2024.103107
DO - 10.1016/j.aei.2024.103107
M3 - 文章
AN - SCOPUS:85214567426
SN - 1474-0346
VL - 65
JO - Advanced Engineering Informatics
JF - Advanced Engineering Informatics
M1 - 103107
ER -