跳到主要导航 跳到搜索 跳到主要内容

Improving Software Defect Prediction in Noisy Imbalanced Datasets

  • Haoxiang Shi
  • , Jun Ai
  • , Jingyu Liu*
  • , Jiaxi Xu
  • *此作品的通讯作者
  • Beihang University
  • China Electronic Product Reliability and Environmental Testing Research Institute

科研成果: 期刊稿件文章同行评审

摘要

Software defect prediction is a popular method for optimizing software testing and improving software quality and reliability. However, software defect datasets usually have quality problems, such as class imbalance and data noise. Oversampling by generating the minority class samples is one of the most well-known methods to improving the quality of datasets; however, it often introduces overfitting noise to datasets. To better improve the quality of these datasets, this paper proposes a method called US-PONR, which uses undersampling to remove duplicate samples from version iterations and then uses oversampling through propensity score matching to reduce class imbalance and noise samples in datasets. The effectiveness of this method was validated in a software prediction experiment that involved 24 versions of software data in 11 projects from PROMISE in noisy environments that varied from 0% to 30% noise level. The experiments showed a significant improvement in the quality of datasets pre-processed by US-PONR in noisy imbalanced datasets, especially the noisiest ones, compared with 12 other advanced dataset processing methods. The experiments also demonstrated that the US-PONR method can effectively identify the label noise samples and remove them.

源语言英语
文章编号10466
期刊Applied Sciences (Switzerland)
13
18
DOI
出版状态已出版 - 9月 2023

学术指纹

探究 'Improving Software Defect Prediction in Noisy Imbalanced Datasets' 的科研主题。它们共同构成独一无二的学术指纹。

引用此