Abstract
Online streaming feature selection plays a critical role in reducing the computing costs and time required for analyzing high-dimensional data streams. In practical big data applications, streaming features commonly contain massive missing data due to unpredictable factors such as instrument failures or privacy issues. However, most existing approaches are designed to work only on complete data, and previous methods utilizing latent factor analysis required manual parameter searches and were restricted to specific data types. To address these limitations, this paper proposes a novel Online Sparse Streaming Feature Selection via Gaussian Copula. The primary methodology involves employing the Gaussian copula to estimate missing values in sparse streaming features, while integrating the expectation maximization algorithm to adaptively update model parameters during the selection process. This design eliminates the need for manual optimal parameter searches and enhances the applicability of the model in real-world online scenarios. Extensive experiments conducted on twelve datasets demonstrate that the proposed algorithm is superior to eight related state-of-the-art methods in addressing sparse streaming feature selection problems. The source code of OS2FSG algorithm is available at: https://anonymous.4open.science/r/OS2FSG-26Y3M17D/.
| Original language | English |
|---|---|
| Article number | 134273 |
| Journal | Neurocomputing |
| Volume | 698 |
| DOIs | |
| State | Published - 14 Oct 2026 |
Keywords
- Feature Selection
- Gaussian Copula
- Machine Learning
- Missing Data
- Sparse Streaming Feature Selection
- Streaming Data
Fingerprint
Dive into the research topics of 'OS2FSG: Online sparse streaming feature selection via Gaussian copula'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver