TY - GEN
T1 - From Human Oral Instructions to General Representations of Knowledge
T2 - 6th International Conference on Cognitive Systems and Signal Processing, ICCSIP 2021
AU - Chen, Shiyu
AU - Zhao, Yongjia
AU - Lei, Xiaoyong
AU - Qi, Tao
AU - Liu, Kan
N1 - Publisher Copyright:
© 2022, Springer Nature Singapore Pte Ltd.
PY - 2022
Y1 - 2022
N2 - Converting human oral instructions into general representations of knowledge which robots can understand, can realize more advanced behaviors for unmanned driving, drones, robots and other fields. This paper presents a novel paradigm based on two-stage structure that transforms human oral instructions into general representations of operating knowledge. Firstly, a Speech-to-Text module is used to realize automatic speech recognition, which results in a text expression for the input oral instruction. In the second stage, a Text-to-Knowledge module is used to realize natural language understanding, which introduces and visualizes task-related knowledge graphs converted from the text expression. To validate this paradigm, the task of computer motherboard assembly was chosen as an example, and a low-cost monophonic speech corpus named PC-CORPUS was built. This PC-CORPUS, is 3 h 44 min long and has 2278 wave audios. 14 speakers from different accent areas in China were invited in the recording. Then experiments were designed and carried out on this paradigm with PC-CORPUS. The average edit distance of this paradigm is 10.87%, comparing the triples of final visual knowledge representations with the ones were labeled by the experts.
AB - Converting human oral instructions into general representations of knowledge which robots can understand, can realize more advanced behaviors for unmanned driving, drones, robots and other fields. This paper presents a novel paradigm based on two-stage structure that transforms human oral instructions into general representations of operating knowledge. Firstly, a Speech-to-Text module is used to realize automatic speech recognition, which results in a text expression for the input oral instruction. In the second stage, a Text-to-Knowledge module is used to realize natural language understanding, which introduces and visualizes task-related knowledge graphs converted from the text expression. To validate this paradigm, the task of computer motherboard assembly was chosen as an example, and a low-cost monophonic speech corpus named PC-CORPUS was built. This PC-CORPUS, is 3 h 44 min long and has 2278 wave audios. 14 speakers from different accent areas in China were invited in the recording. Then experiments were designed and carried out on this paradigm with PC-CORPUS. The average edit distance of this paradigm is 10.87%, comparing the triples of final visual knowledge representations with the ones were labeled by the experts.
KW - Computer assembly
KW - Knowledge graph
KW - Natural language understanding
UR - https://www.scopus.com/pages/publications/85123586102
U2 - 10.1007/978-981-16-9247-5_29
DO - 10.1007/978-981-16-9247-5_29
M3 - 会议稿件
AN - SCOPUS:85123586102
SN - 9789811692468
T3 - Communications in Computer and Information Science
SP - 374
EP - 388
BT - Cognitive Systems and Information Processing - 6th International Conference, ICCSIP 2021, Revised Selected Papers
A2 - Sun, Fuchun
A2 - Hu, Dewen
A2 - Wermter, Stefan
A2 - Yang, Lei
A2 - Liu, Huaping
A2 - Fang, Bin
PB - Springer Science and Business Media Deutschland GmbH
Y2 - 20 November 2021 through 21 November 2021
ER -