Skip to main navigation Skip to search Skip to main content

Efficient Parallel Auto-encoder Based on Spark

  • CAS - Institute of Computing Technology
  • Yanshan University

Research output: Contribution to journalArticlepeer-review

Abstract

How to find the good representation from raw data is a key and very important issue in machine learning. Most traditional approaches are based on the relationship among data or utilize simple linear combination, in which deep learning algorithm can perform very well in various machine learning tasks and achieve very good representations. However, most existing algorithms are implemented in serial, which cannot handle large-scale data. This paper proposes an effective parallel auto-encoder (PAE) based on Spark. The proposed PAE not only can learn satisfying representation, but also can speed up the executing time based on Spark. And then the paper adapts PAE to deal with the sparse data. Experiments conducted on two tasks, i.e., classification and collaborative filtering, demonstrate the effectiveness and efficiency of the proposed PAE.

Original languageEnglish
Pages (from-to)65-74
Number of pages10
JournalShuju Caiji Yu Chuli/Journal of Data Acquisition and Processing
Volume33
Issue number1
DOIs
StatePublished - Jan 2018
Externally publishedYes

Keywords

  • Auto-encoder
  • Deep learning
  • Feature learning
  • Machine learning
  • Spark

Fingerprint

Dive into the research topics of 'Efficient Parallel Auto-encoder Based on Spark'. Together they form a unique fingerprint.

Cite this