Skip to main navigation Skip to search Skip to main content

ACID: Association Correction for Imbalanced Data in GWAS

  • Tsinghua University
  • University of California at San Francisco

Research output: Contribution to journalArticlepeer-review

Abstract

Genome-wide association study (GWAS) has been widely witnessed as a powerful tool for revealing suspicious loci from various diseases. However, real world GWAS tasks always suffer from the data imbalance problem of sufficient control samples and limited case samples. This imbalance issue can cause serious biases to the result and thus leads to losses of significance for true causal markers. To tackle this problem, we proposed a computational framework to perform association correction for imbalanced data (ACID) that could potentially improve the performance of GWAS under the imbalance condition. ACID is inspired by the imbalance learning theory but is particularly modified to address the task of association discovery from sequential genomic data. Simulation studies demonstrate ACID can dramatically improve the power of traditional GWAS method on the dataset with severe imbalances. We further applied ACID to two imbalanced datasets (gastric cancer and bladder cancer) to conduct genome wide association analysis. Experimental results indicate that our method has better abilities in identifying suspicious loci than the regression approach and shows consistencies with existing discoveries.

Original languageEnglish
Article number7565545
Pages (from-to)316-322
Number of pages7
JournalIEEE/ACM Transactions on Computational Biology and Bioinformatics
Volume15
Issue number1
DOIs
StatePublished - 1 Jan 2018
Externally publishedYes

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 3 - Good Health and Well-being
    SDG 3 Good Health and Well-being

Keywords

  • Genome-wide association study
  • genome structures
  • hidden Markov models
  • imbalance learning

Fingerprint

Dive into the research topics of 'ACID: Association Correction for Imbalanced Data in GWAS'. Together they form a unique fingerprint.

Cite this