TY - JOUR
T1 - CPCA
T2 - A feature semantics based crowd dimension reduction framework
AU - Zhang, Yuanyuan
AU - Gao, Dawei
AU - Luo, Jie
AU - Xu, Ke
N1 - Publisher Copyright:
© 2018 IEEE.
PY - 2018
Y1 - 2018
N2 - Dimension reduction plays an important role in practical big data analysis and data mining applications. However, popular dimension reduction techniques, such as principal component analysis (PCA), are known to be computation-intensive and are considered as a computation bottleneck for data processing and mining. In this paper, we propose to reduce the computation of PCA via crowdsourcing, a paradigm that accomplishes hard-to-compute problems leveraging collective intelligence. We design CPCA, crowd principal component analysis, a novel crowd-based dimension reduction framework. The CPCA designs tasks for crowd workers to obtain the relations among features based on their semantics and formulates a weighted graph from the collected answers to derive the covariance matrix and the principal components. We prove the correctness of CPCA and conduct extensive evaluations on real datasets. Experimental results show that CPCA could achieve significantly reduction on the computational cost in terms of both time and memory, which lowers the bar for learning.
AB - Dimension reduction plays an important role in practical big data analysis and data mining applications. However, popular dimension reduction techniques, such as principal component analysis (PCA), are known to be computation-intensive and are considered as a computation bottleneck for data processing and mining. In this paper, we propose to reduce the computation of PCA via crowdsourcing, a paradigm that accomplishes hard-to-compute problems leveraging collective intelligence. We design CPCA, crowd principal component analysis, a novel crowd-based dimension reduction framework. The CPCA designs tasks for crowd workers to obtain the relations among features based on their semantics and formulates a weighted graph from the collected answers to derive the covariance matrix and the principal components. We prove the correctness of CPCA and conduct extensive evaluations on real datasets. Experimental results show that CPCA could achieve significantly reduction on the computational cost in terms of both time and memory, which lowers the bar for learning.
KW - Dimensionality reduction
KW - crowdsourcing
KW - machine learning
KW - principal component analysis
UR - https://www.scopus.com/pages/publications/85056193955
U2 - 10.1109/ACCESS.2018.2879011
DO - 10.1109/ACCESS.2018.2879011
M3 - 文章
AN - SCOPUS:85056193955
SN - 2169-3536
VL - 6
SP - 73191
EP - 73199
JO - IEEE Access
JF - IEEE Access
M1 - 8519735
ER -