BioRel: A Large-Scale Dataset for Biomedical Relation Extraction

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Valuable biomedical knowledge usually exists in the form of electronic publications and literature, which is growing at an enormous rate. Relation extraction plays a critical role in discovering such knowledge and transform them into structural form. Previous relation extraction datasets in biomedical domain are mainly human-annotated, whose scales are usually limited due to their labor-intensive and time-consuming nature. In this paper, we present BioRel, a large-scale dataset constructed by using Unified Medical Language System (UMLS) as knowledge base and Medline as corpus. Entities in sentences of Medline are identified and linked to UMLS by Metamap. Relation label for each sentence is recognized using distant supervision. We adapt both state-of-the-art deep learning and statistical machine learning methods as baseline models and conduct comprehensive experiments on BioRel. Experimental results show that BioRel is suitable for training and evaluating relation extraction models for both deep learning and statistical methods by providing both reasonable baseline performance and many remaining challenges.

Original languageEnglish
Title of host publicationProceedings - 2019 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2019
EditorsIllhoi Yoo, Jinbo Bi, Xiaohua Tony Hu
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages1801-1808
Number of pages8
ISBN (Electronic)9781728118673
DOIs
StatePublished - Nov 2019
Event2019 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2019 - San Diego, United States
Duration: 18 Nov 201921 Nov 2019

Publication series

NameProceedings - 2019 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2019

Conference

Conference2019 IEEE International Conference on Bioinformatics and Biomedicine, BIBM 2019
Country/TerritoryUnited States
CitySan Diego
Period18/11/1921/11/19

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 3 - Good Health and Well-being
    SDG 3 Good Health and Well-being

Keywords

  • Medline
  • dataset
  • distant supervision
  • relation extraction

Fingerprint

Dive into the research topics of 'BioRel: A Large-Scale Dataset for Biomedical Relation Extraction'. Together they form a unique fingerprint.

Cite this