摘要
How to find the good representation from raw data is a key and very important issue in machine learning. Most traditional approaches are based on the relationship among data or utilize simple linear combination, in which deep learning algorithm can perform very well in various machine learning tasks and achieve very good representations. However, most existing algorithms are implemented in serial, which cannot handle large-scale data. This paper proposes an effective parallel auto-encoder (PAE) based on Spark. The proposed PAE not only can learn satisfying representation, but also can speed up the executing time based on Spark. And then the paper adapts PAE to deal with the sparse data. Experiments conducted on two tasks, i.e., classification and collaborative filtering, demonstrate the effectiveness and efficiency of the proposed PAE.
| 源语言 | 英语 |
|---|---|
| 页(从-至) | 65-74 |
| 页数 | 10 |
| 期刊 | Shuju Caiji Yu Chuli/Journal of Data Acquisition and Processing |
| 卷 | 33 |
| 期 | 1 |
| DOI | |
| 出版状态 | 已出版 - 1月 2018 |
| 已对外发布 | 是 |
学术指纹
探究 'Efficient Parallel Auto-encoder Based on Spark' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver