跳到主要导航 跳到搜索 跳到主要内容

Pretraining Language Models with Text-Attributed Heterogeneous Graphs

  • Beihang University
  • Zhongguancun Laboratory

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

In many real-world scenarios (e.g., academic networks, social platforms), different types of entities are not only associated with texts but also connected by various relationships, which can be abstracted as Text-Attributed Heterogeneous Graphs (TAHGs). Current pretraining tasks for Language Models (LMs) primarily focus on separately learning the textual information of each entity and overlook the crucial aspect of capturing topological connections among entities in TAHGs. In this paper, we present a new pretraining framework for LMs that explicitly considers the topological and heterogeneous information in TAHGs. Firstly, we define a context graph as neighborhoods of a target node within specific orders and propose a topology-aware pretraining task to predict nodes involved in the context graph by jointly optimizing an LM and an auxiliary heterogeneous graph neural network. Secondly, based on the observation that some nodes are text-rich while others have little text, we devise a text augmentation strategy to enrich textless nodes with their neighbors' texts for handling the imbalance issue. We conduct link prediction and node classification tasks on three datasets from various domains. Experimental results demonstrate the superiority of our approach over existing methods and the rationality of each design. Our code is available at https://github.com/Hope-Rita/THLM.

源语言英语
主期刊名Findings of the Association for Computational Linguistics
主期刊副标题EMNLP 2023
出版商Association for Computational Linguistics (ACL)
10316-10333
页数18
ISBN(电子版)9798891760615
DOI
出版状态已出版 - 2023
活动2023 Findings of the Association for Computational Linguistics: EMNLP 2023 - Hybrid, 新加坡
期限: 6 12月 202310 12月 2023

出版系列

姓名Findings of the Association for Computational Linguistics: EMNLP 2023

会议

会议2023 Findings of the Association for Computational Linguistics: EMNLP 2023
国家/地区新加坡
Hybrid
时期6/12/2310/12/23

指纹

探究 'Pretraining Language Models with Text-Attributed Heterogeneous Graphs' 的科研主题。它们共同构成独一无二的指纹。

引用此