TY - JOUR
T1 - Scalable Database-Driven KGs can help Text-to-SQL
AU - Li, Zhongqiu
AU - Wu, Zhenhe
AU - Li, Mengxiang
AU - He, Zhongjiang
AU - Fang, Ruiyu
AU - Zhang, Jie
AU - Zhao, Yu
AU - Li, Yongxiang
AU - Li, Zhoujun
AU - Song, Shuangyong
N1 - Publisher Copyright:
© 2024 Copyright for this paper by its authors.
PY - 2024
Y1 - 2024
N2 - The text-to-SQL task aims to covert natural language questions into SQL queries. Large Language Models (LLMs) have demonstrated remarkable performance on this task, which relied on in-context learing or Supervised Fine-Tuning (SFT). However, the heterogeneity of database and the complexity of the knowledge acquisition process pose significant challenges in previous works. To address these, we propose a novel text-to-SQL framework that enhances the performance of LLMs through Knowledge Graphs (KGs). We construct the KGs based on schemas, which are structured representations of the relationships and attributes within the databases. Then, we utilize LLMs to extract descriptions and dependencies from historical queries, which are used to complete contextual knowledge in KGs. We leverage retrieval model to recall benefit nodes and edges from KGs and then employ LLMs to generate task-specific evidence. Based on the evidence and retrieved information, we define a unified KGs-based schema for LLMs to generate SQL queries. Our paper conducts experiments on public datasets BIRD and Spider, and the results indicate that our framework significantly improves the text-to-SQL performance.
AB - The text-to-SQL task aims to covert natural language questions into SQL queries. Large Language Models (LLMs) have demonstrated remarkable performance on this task, which relied on in-context learing or Supervised Fine-Tuning (SFT). However, the heterogeneity of database and the complexity of the knowledge acquisition process pose significant challenges in previous works. To address these, we propose a novel text-to-SQL framework that enhances the performance of LLMs through Knowledge Graphs (KGs). We construct the KGs based on schemas, which are structured representations of the relationships and attributes within the databases. Then, we utilize LLMs to extract descriptions and dependencies from historical queries, which are used to complete contextual knowledge in KGs. We leverage retrieval model to recall benefit nodes and edges from KGs and then employ LLMs to generate task-specific evidence. Based on the evidence and retrieved information, we define a unified KGs-based schema for LLMs to generate SQL queries. Our paper conducts experiments on public datasets BIRD and Spider, and the results indicate that our framework significantly improves the text-to-SQL performance.
KW - Knowledge Generation
KW - Knowledge Graph
KW - Large Language Models
KW - Text-to-SQL
UR - https://www.scopus.com/pages/publications/85210231219
M3 - 会议文章
AN - SCOPUS:85210231219
SN - 1613-0073
VL - 3828
JO - CEUR Workshop Proceedings
JF - CEUR Workshop Proceedings
T2 - ISWC 2024 Posters, Demos and Industry Tracks: From Novel Ideas to Industrial Practice, ISWC-Posters-Demos-Industry 2024
Y2 - 11 November 2024 through 15 November 2024
ER -