TY - GEN
T1 - Taking too many destinations can be bad for traceroute sampling
AU - Qiao, Zhongliang
AU - Chen, Mingming
AU - Xu, Ke
PY - 2010
Y1 - 2010
N2 - Considerable effort has been spent on collecting all the information of routers by using traceroute-like probes in the router-level topology measurements. This method has been argued to introduce uncontrolled sampling biases on statistical properties of the sample graph and heavy load to the network being measured. In order to improve the quality of the maps induced by the method, researchers are starting to investigate the deployment of large-scale distributed systems. But the lack of sources, the additional load introduced to the network and the potential uncontrolled scale in the IPv6 network cast a shadow over this direction. In this paper, we study traceroute sampling to represent the topology. Instead of finding a general strategy that would match all the graph properties, we focus on testing the impact of the proportion of destinations and sources on a single or several properties of the graph. We argue that, in order to obtain a more accurate sample graph, the general method of taking a small set of sources to perform traceroute-like probes to a large set of destinations is very unreasonable. Our results obtained from simulated experiments show that as the proportion of destinations and sources increases, the resulting properties of the sample graph can differ sharply from the underlying graph. The results also show that there is no single perfect proportion answer to meet all the graph properties but the small often perform better than the large overall. When we do the same measurement on several real-world networks, we find strong evidence for sampling bias because of taking so many destinations.
AB - Considerable effort has been spent on collecting all the information of routers by using traceroute-like probes in the router-level topology measurements. This method has been argued to introduce uncontrolled sampling biases on statistical properties of the sample graph and heavy load to the network being measured. In order to improve the quality of the maps induced by the method, researchers are starting to investigate the deployment of large-scale distributed systems. But the lack of sources, the additional load introduced to the network and the potential uncontrolled scale in the IPv6 network cast a shadow over this direction. In this paper, we study traceroute sampling to represent the topology. Instead of finding a general strategy that would match all the graph properties, we focus on testing the impact of the proportion of destinations and sources on a single or several properties of the graph. We argue that, in order to obtain a more accurate sample graph, the general method of taking a small set of sources to perform traceroute-like probes to a large set of destinations is very unreasonable. Our results obtained from simulated experiments show that as the proportion of destinations and sources increases, the resulting properties of the sample graph can differ sharply from the underlying graph. The results also show that there is no single perfect proportion answer to meet all the graph properties but the small often perform better than the large overall. When we do the same measurement on several real-world networks, we find strong evidence for sampling bias because of taking so many destinations.
KW - Graph sampling
KW - Network topology
KW - Traceroute
UR - https://www.scopus.com/pages/publications/77958582024
U2 - 10.1109/ICCCN.2010.5560162
DO - 10.1109/ICCCN.2010.5560162
M3 - 会议稿件
AN - SCOPUS:77958582024
SN - 9781424471164
T3 - Proceedings - International Conference on Computer Communications and Networks, ICCCN
BT - 2010 Proceedings of 19th International Conference on Computer Communications and Networks, ICCCN 2010
T2 - 2010 19th International Conference on Computer Communications and Networks, ICCCN 2010
Y2 - 2 August 2010 through 5 August 2010
ER -