TY - GEN
T1 - ManufVisSGG
T2 - 20th IEEE International Conference on Automation Science and Engineering, CASE 2024
AU - Yan, Zhijie
AU - Wang, Zuoxu
AU - Li, Shufei
AU - Li, Mingrui
AU - Liang, Xinxin
AU - Liu, Jihong
N1 - Publisher Copyright:
© 2024 IEEE.
PY - 2024
Y1 - 2024
N2 - To establish a cognitive manufacturing system, scene graph generation (SGG) that lets machines/robots understand objects and their relations under varied scenarios is an essential task. Existing research on SGG primarily focuses on detection and panoptic segmentation approaches, where objects are identified through bounding boxes or panoptic segmentation, followed by the prediction of their pairwise relationships. This process means that the quality of the final scene graph predictions is heavily influenced by the quality of costly annotations. To tackle this issue, we propose the Manufacturing Visual Scene Graph Generation (ManufVisSGG) method, a simple yet powerful approach that leverages the capabilities of Vision-Language Models (VLMs) to generate scene graphs quickly and accurately without any additional object annotations. Furthermore, leveraging the ManufVisSGG method, we have implemented a meticulous annotation procedure to compile a high-quality manufacturing scene graph generation (MSG) dataset, comprising 10,000 images of manufacturing and other industrial scenes. Through comparisons with various scene graph generation methods and benchmarks across two other datasets, we have showcased the superiority of the ManufVisSGG method and underscored the benefits of the MSG dataset over existing datasets.
AB - To establish a cognitive manufacturing system, scene graph generation (SGG) that lets machines/robots understand objects and their relations under varied scenarios is an essential task. Existing research on SGG primarily focuses on detection and panoptic segmentation approaches, where objects are identified through bounding boxes or panoptic segmentation, followed by the prediction of their pairwise relationships. This process means that the quality of the final scene graph predictions is heavily influenced by the quality of costly annotations. To tackle this issue, we propose the Manufacturing Visual Scene Graph Generation (ManufVisSGG) method, a simple yet powerful approach that leverages the capabilities of Vision-Language Models (VLMs) to generate scene graphs quickly and accurately without any additional object annotations. Furthermore, leveraging the ManufVisSGG method, we have implemented a meticulous annotation procedure to compile a high-quality manufacturing scene graph generation (MSG) dataset, comprising 10,000 images of manufacturing and other industrial scenes. Through comparisons with various scene graph generation methods and benchmarks across two other datasets, we have showcased the superiority of the ManufVisSGG method and underscored the benefits of the MSG dataset over existing datasets.
UR - https://www.scopus.com/pages/publications/85202074549
U2 - 10.1109/CASE59546.2024.10711649
DO - 10.1109/CASE59546.2024.10711649
M3 - 会议稿件
AN - SCOPUS:85202074549
T3 - IEEE International Conference on Automation Science and Engineering
SP - 1632
EP - 1637
BT - 2024 IEEE 20th International Conference on Automation Science and Engineering, CASE 2024
PB - IEEE Computer Society
Y2 - 28 August 2024 through 1 September 2024
ER -