跳到主要导航 跳到搜索 跳到主要内容

Efficient Mixture of Experts based on Large Language Models for Low-Resource Data Preprocessing

  • Mengyi Yan
  • , Yaoshu Wang*
  • , Kehan Pang
  • , Min Xie
  • , Jianxin Li
  • *此作品的通讯作者
  • Beihang University
  • Shenzhen Institute of Computing Sciences

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Data preprocessing (DP) that transforms erroneous and raw data to a clean version is a cornerstone of the data mining pipeline. Due to the diverse requirements of downstream tasks, data scientists and domain experts have to handcraft domain-specific rules or train ML models with annotated examples, which is costly/time-consuming. In this paper, we present MELD (<u>Mixture of <u>Experts on <u>Large Language Models for <u>Data Preprocessing), a universal solver for low-resource DP. MELD adopts a Mixture-of-Experts (MoE) architecture that enables the amalgamation and enhancement of domain-specific experts trained on limited annotated examples. To fine-tune MELD, we develop a suite of expert-tuning and MoE-tuning techniques, including a retrieval augmented generation (RAG) system, meta-path search for data augmentation, expert refinement and router network training based on information bottleneck. To further verify the effectiveness of MELD, we theoretically prove that MoE in MELD is superior than a single expert and the router network is able to dispatch data to the right experts. Finally, we conducted extensive experiments on 19 datasets over 10 DP tasks to show that MELD outperforms the state-of-the-art methods in both effectiveness and efficiency. More importantly, MELD is able to be fine-tuned in a low-resource environment, e.g. a local, single and low-priced 3090 GPU.

源语言英语
主期刊名KDD 2024 - Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining
出版商Association for Computing Machinery
3690-3701
页数12
ISBN(电子版)9798400704901
DOI
出版状态已出版 - 24 8月 2024
活动30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2024 - Barcelona, 西班牙
期限: 25 8月 202429 8月 2024

出版系列

姓名Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
ISSN(印刷版)2154-817X

会议

会议30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2024
国家/地区西班牙
Barcelona
时期25/08/2429/08/24

指纹

探究 'Efficient Mixture of Experts based on Large Language Models for Low-Resource Data Preprocessing' 的科研主题。它们共同构成独一无二的指纹。

引用此