Abstract
Genome-wide association study (GWAS) has been widely witnessed as a powerful tool for revealing suspicious loci from various diseases. However, real world GWAS tasks always suffer from the data imbalance problem of sufficient control samples and limited case samples. This imbalance issue can cause serious biases to the result and thus leads to losses of significance for true causal markers. To tackle this problem, we proposed a computational framework to perform association correction for imbalanced data (ACID) that could potentially improve the performance of GWAS under the imbalance condition. ACID is inspired by the imbalance learning theory but is particularly modified to address the task of association discovery from sequential genomic data. Simulation studies demonstrate ACID can dramatically improve the power of traditional GWAS method on the dataset with severe imbalances. We further applied ACID to two imbalanced datasets (gastric cancer and bladder cancer) to conduct genome wide association analysis. Experimental results indicate that our method has better abilities in identifying suspicious loci than the regression approach and shows consistencies with existing discoveries.
| Original language | English |
|---|---|
| Article number | 7565545 |
| Pages (from-to) | 316-322 |
| Number of pages | 7 |
| Journal | IEEE/ACM Transactions on Computational Biology and Bioinformatics |
| Volume | 15 |
| Issue number | 1 |
| DOIs | |
| State | Published - 1 Jan 2018 |
| Externally published | Yes |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 3 Good Health and Well-being
Keywords
- Genome-wide association study
- genome structures
- hidden Markov models
- imbalance learning
Fingerprint
Dive into the research topics of 'ACID: Association Correction for Imbalanced Data in GWAS'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver