跳到主要导航 跳到搜索 跳到主要内容

Easy Dataset: A Unified and Extensible Framework for Synthesizing LLM Fine-Tuning Data from Unstructured Documents

  • Beihang University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Large language models (LLMs) have shown impressive performance on general-purpose tasks, yet adapting them to specific domains remains challenging due to the scarcity of high-quality domain data. Existing data synthesis tools often struggle to extract reliable fine-tuning data from heterogeneous documents effectively. To address this limitation, we propose Easy Dataset, a unified framework for synthesizing fine-tuning data from unstructured documents via an intuitive graphical user interface (GUI). Specifically, Easy Dataset allows users to easily configure text extraction models and chunking strategies to transform raw documents into coherent text chunks. It then leverages a persona-driven prompting approach to generate diverse question-answer pairs using public-available LLMs. Throughout the pipeline, a human-in-the-loop visual interface facilitates the review and refinement of intermediate outputs to ensure data quality. Experiments on a financial question-answering task show that fine-tuning LLMs on the synthesized dataset significantly improves domain-specific performance while preserving general knowledge.

源语言英语
主期刊名EMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the System Demonstrations
编辑Ivan Habernal, Peter Schulam, Jorg Tiedemann
出版商Association for Computational Linguistics (ACL)
960-968
页数9
ISBN(电子版)9798891763340
DOI
出版状态已出版 - 2025
活动2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, EMNLP 2025 - Suzhou, 中国
期限: 4 11月 20259 11月 2025

出版系列

姓名EMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the System Demonstrations

会议

会议2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, EMNLP 2025
国家/地区中国
Suzhou
时期4/11/259/11/25

指纹

探究 'Easy Dataset: A Unified and Extensible Framework for Synthesizing LLM Fine-Tuning Data from Unstructured Documents' 的科研主题。它们共同构成独一无二的指纹。

引用此