TY - GEN
T1 - Using clustering and transitivity to reduce the costs of crowdsourced entity resolution
AU - Guo, Lisha
AU - Sun, Hailong
AU - Liu, Xudong
N1 - Publisher Copyright:
Copyright 2014 ACM.
PY - 2014/11/17
Y1 - 2014/11/17
N2 - Entity resolution is the process of identifying the data records representing the same entity. ER is a highly important problem in software and application domains. For example, detecting duplicate bug reports with ER can greatly save developing efforts. In most cases, humans can perform better than computer algorithms due to complex semantic analysis involved in ER. In light of this, crowdsourcing has been successfully incorporated into ER to improve its accuracy. However, compared with computer methods, crowdsourcing is subject to higher costs. In this work, we propose a method to reduce the number of questions asked to people with clustering and transitivity analysis. Firstly, with appropriate choosing of two similarity thresholds, we use unsupervised machine learning to cluster records into multiple clusters on the basis of certain similarity metrics. In this way, we prune away the record pairs with no need for asking people. Secondly, we design a cluster merging algorithm with efficient selection of crowdsourced questions and leveraging data transitivity to detect the across-cluster records corresponding to the same entity. Finally, we conduct extensive experiments with two real-world datasets and the results show our method significantly outperform existing methods in terms of incurred costs and the F1 metric.
AB - Entity resolution is the process of identifying the data records representing the same entity. ER is a highly important problem in software and application domains. For example, detecting duplicate bug reports with ER can greatly save developing efforts. In most cases, humans can perform better than computer algorithms due to complex semantic analysis involved in ER. In light of this, crowdsourcing has been successfully incorporated into ER to improve its accuracy. However, compared with computer methods, crowdsourcing is subject to higher costs. In this work, we propose a method to reduce the number of questions asked to people with clustering and transitivity analysis. Firstly, with appropriate choosing of two similarity thresholds, we use unsupervised machine learning to cluster records into multiple clusters on the basis of certain similarity metrics. In this way, we prune away the record pairs with no need for asking people. Secondly, we design a cluster merging algorithm with efficient selection of crowdsourced questions and leveraging data transitivity to detect the across-cluster records corresponding to the same entity. Finally, we conduct extensive experiments with two real-world datasets and the results show our method significantly outperform existing methods in terms of incurred costs and the F1 metric.
KW - Clustering
KW - Crowdsourcing
KW - Entity resolution
KW - Software
KW - Transitivity
UR - https://www.scopus.com/pages/publications/84945406551
U2 - 10.1145/2666539.2666568
DO - 10.1145/2666539.2666568
M3 - 会议稿件
AN - SCOPUS:84945406551
T3 - 1st International Workshop on Crowd-Based Software Development Methods and Technologies, CrowdSoft 2014 - Proceedings
SP - 13
EP - 18
BT - 1st International Workshop on Crowd-Based Software Development Methods and Technologies, CrowdSoft 2014 - Proceedings
PB - Association for Computing Machinery
T2 - 1st International Workshop on Crowd-Based Software Development Methods and Technologies, CrowdSoft 2014
Y2 - 17 November 2014
ER -