TY - GEN
T1 - DisITQ
T2 - 9th International Symposium on Computational Intelligence and Design, ISCID 2016
AU - Chen, Qun
AU - Lang, Bo
AU - Liu, Xianglong
AU - Gu, Zepeng
N1 - Publisher Copyright:
© 2016 IEEE.
PY - 2016/7/2
Y1 - 2016/7/2
N2 - In the field of big data retrieval, hashing based approxi-mate nearest neighbors (ANN) search has attracted many at-tentions. However, most existing hashing algorithms are learned from the centralized settings and based on small scale datasets, or in other words, they are single machine ap-proaches which load the training data into memory to get models. For big data processing, models learned from large scale datasets which have the properties of big data such as variety often have better performance. However, there are two critical problems when training datasets are in very large size. First, a single compute node can't load all the data into memory to train hashing models. Second, in real-word appli-cations, the data is often stored or even collected in a distrib-uted manner, and it's infeasible to gather all data into a fu-sion center because of the prohibitively expensive communi-cation and computation overhead. In this article, we present a distributed learning algorithm which is based on MapReduce and Iterative Quantization (ITQ) to train hashing functions. The proposed method, named as distributed iterative quanti-zation hashing (DisITQ), can not only be performed on large scale datasets, but can also be applied to distributed data storing scenarios. Massive experiments carried out on large scale datasets demonstrate the time efficiency and the accu-racy advantages of the method we proposed in comparison with the state-of-the-art hashing algorithms.
AB - In the field of big data retrieval, hashing based approxi-mate nearest neighbors (ANN) search has attracted many at-tentions. However, most existing hashing algorithms are learned from the centralized settings and based on small scale datasets, or in other words, they are single machine ap-proaches which load the training data into memory to get models. For big data processing, models learned from large scale datasets which have the properties of big data such as variety often have better performance. However, there are two critical problems when training datasets are in very large size. First, a single compute node can't load all the data into memory to train hashing models. Second, in real-word appli-cations, the data is often stored or even collected in a distrib-uted manner, and it's infeasible to gather all data into a fu-sion center because of the prohibitively expensive communi-cation and computation overhead. In this article, we present a distributed learning algorithm which is based on MapReduce and Iterative Quantization (ITQ) to train hashing functions. The proposed method, named as distributed iterative quanti-zation hashing (DisITQ), can not only be performed on large scale datasets, but can also be applied to distributed data storing scenarios. Massive experiments carried out on large scale datasets demonstrate the time efficiency and the accu-racy advantages of the method we proposed in comparison with the state-of-the-art hashing algorithms.
KW - Big data
KW - Distributed hashing
KW - MapReduce
UR - https://www.scopus.com/pages/publications/85013683064
U2 - 10.1109/ISCID.2016.2036
DO - 10.1109/ISCID.2016.2036
M3 - 会议稿件
AN - SCOPUS:85013683064
T3 - Proceedings - 2016 9th International Symposium on Computational Intelligence and Design, ISCID 2016
SP - 118
EP - 123
BT - Proceedings - 2016 9th International Symposium on Computational Intelligence and Design, ISCID 2016
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 10 December 2016 through 11 December 2016
ER -