Skip to main navigation Skip to search Skip to main content

Scalable Database-Driven KGs can help Text-to-SQL

  • Zhongqiu Li
  • , Zhenhe Wu
  • , Mengxiang Li
  • , Zhongjiang He
  • , Ruiyu Fang
  • , Jie Zhang
  • , Yu Zhao
  • , Yongxiang Li
  • , Zhoujun Li*
  • , Shuangyong Song*
  • *Corresponding author for this work
  • China Telecommunications
  • Beihang University

Research output: Contribution to journalConference articlepeer-review

Abstract

The text-to-SQL task aims to covert natural language questions into SQL queries. Large Language Models (LLMs) have demonstrated remarkable performance on this task, which relied on in-context learing or Supervised Fine-Tuning (SFT). However, the heterogeneity of database and the complexity of the knowledge acquisition process pose significant challenges in previous works. To address these, we propose a novel text-to-SQL framework that enhances the performance of LLMs through Knowledge Graphs (KGs). We construct the KGs based on schemas, which are structured representations of the relationships and attributes within the databases. Then, we utilize LLMs to extract descriptions and dependencies from historical queries, which are used to complete contextual knowledge in KGs. We leverage retrieval model to recall benefit nodes and edges from KGs and then employ LLMs to generate task-specific evidence. Based on the evidence and retrieved information, we define a unified KGs-based schema for LLMs to generate SQL queries. Our paper conducts experiments on public datasets BIRD and Spider, and the results indicate that our framework significantly improves the text-to-SQL performance.

Keywords

  • Knowledge Generation
  • Knowledge Graph
  • Large Language Models
  • Text-to-SQL

Fingerprint

Dive into the research topics of 'Scalable Database-Driven KGs can help Text-to-SQL'. Together they form a unique fingerprint.

Cite this