TY - GEN
T1 - AnaSearch
T2 - 14th ACM International Conference on Web Search and Data Mining, WSDM 2021
AU - Li, Tongliang
AU - Fang, Lei
AU - Lou, Jian Guang
AU - Li, Zhoujun
AU - Zhang, Dongmei
N1 - Publisher Copyright:
© 2021 Owner/Author.
PY - 2021/8/3
Y1 - 2021/8/3
N2 - Modern search engines retrieve results mainly based on the keyword matching techniques, and thus fail to answer analytical queries like "apps with more than 1 billion monthly active users"or "population growth of the US from 2015 to 2019", which requires numerical reasoning or aggregating results from multiple web pages. Such analytical queries are very common in the data analysis area, the expected results would be structured tables or charts. In most cases, these structured results are not available or accessible, they scatter in various text sources. In this work, we build AnaSearch, a search system to support analytical queries, and return structured results that can be visualized in the form of tables or charts. We collect and build structured quantitative data from the unstructured text on the web automatically. With AnaSearch, data analysts could easily derive insights for decision making with keyword or natural language queries. Specifically, we build AnaSearch under the COVID-19 news data, which makes it easy to compare with manually collected structured data.
AB - Modern search engines retrieve results mainly based on the keyword matching techniques, and thus fail to answer analytical queries like "apps with more than 1 billion monthly active users"or "population growth of the US from 2015 to 2019", which requires numerical reasoning or aggregating results from multiple web pages. Such analytical queries are very common in the data analysis area, the expected results would be structured tables or charts. In most cases, these structured results are not available or accessible, they scatter in various text sources. In this work, we build AnaSearch, a search system to support analytical queries, and return structured results that can be visualized in the form of tables or charts. We collect and build structured quantitative data from the unstructured text on the web automatically. With AnaSearch, data analysts could easily derive insights for decision making with keyword or natural language queries. Specifically, we build AnaSearch under the COVID-19 news data, which makes it easy to compare with manually collected structured data.
KW - data visualization
KW - information retrieval
KW - quantitative information
KW - structured data
UR - https://www.scopus.com/pages/publications/85103044779
U2 - 10.1145/3437963.3441694
DO - 10.1145/3437963.3441694
M3 - 会议稿件
AN - SCOPUS:85103044779
T3 - WSDM 2021 - Proceedings of the 14th ACM International Conference on Web Search and Data Mining
SP - 906
EP - 909
BT - WSDM 2021 - Proceedings of the 14th ACM International Conference on Web Search and Data Mining
PB - Association for Computing Machinery, Inc
Y2 - 8 March 2021 through 12 March 2021
ER -