Skip to main navigation Skip to search Skip to main content

Outlier detection in sparse data with factorization machines

  • Mengxiao Zhu
  • , Charu C. Aggarwal
  • , Shuai Ma*
  • , Hui Zhang
  • , Jinpeng Huai
  • *Corresponding author for this work
  • Beihang University
  • Beijing Advanced Innovation Center for Big Data and Brain Computing
  • IBM

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

In sparse data, a large fraction of the entries take on zero values. Some examples of sparse data include short text snippets (such as tweets in Twitter) or some feature representations of categorical data sets with a large number of values, in which traditional methods for outlier detection typically fail because of the difficulty of computing distances. To address this, it is important to use the latent relations between such values. Factorization machines represent a natural methodology for this, and are naturally designed for the massive-domain setting because of their emphasis on sparse data sets. In this study, we propose an outlier detection approach for sparse data with factorization machines. Factorization machines are also efficient due to their linear complexity in the number of non-zero values. In fact, because of their efficiency, they can even be extended to traditional settings for numerical data by an appropriate feature engineering effort. We show that our approach is both effective and efficient for sparse categorical, short text and numerical data by an extensive experimental study.

Original languageEnglish
Title of host publicationCIKM 2017 - Proceedings of the 2017 ACM Conference on Information and Knowledge Management
PublisherAssociation for Computing Machinery
Pages817-826
Number of pages10
ISBN (Electronic)9781450349185
DOIs
StatePublished - 6 Nov 2017
Event26th ACM International Conference on Information and Knowledge Management, CIKM 2017 - Singapore, Singapore
Duration: 6 Nov 201710 Nov 2017

Publication series

NameInternational Conference on Information and Knowledge Management, Proceedings
VolumePart F131841

Conference

Conference26th ACM International Conference on Information and Knowledge Management, CIKM 2017
Country/TerritorySingapore
CitySingapore
Period6/11/1710/11/17

Keywords

  • Factorization machines
  • Outlier detection
  • Sparse data

Fingerprint

Dive into the research topics of 'Outlier detection in sparse data with factorization machines'. Together they form a unique fingerprint.

Cite this