Skip to main navigation Skip to search Skip to main content

Selecting valuable training samples for SVMs via data structure analysis

  • Defeng Wang*
  • , Lin Shi
  • *Corresponding author for this work
  • Chinese University of Hong Kong

Research output: Contribution to journalArticlepeer-review

Abstract

In spite of its salient properties and wide acceptance, support vector machines (SVMs) still face difficulties in scalability, because solving the quadratic programming (QP) problems in SVMs training is especially costly when dealing with large sets of training data. This paper presents a new algorithm named sample reduction by data structure analysis (SR-DSA) for SVMs to improve their scalability. The SR-DSA utilizes data structure information in selecting the data points valuable in learning the separating plane. As this method is performed completely before SVMs training, it avoids the problem suffered by most sample reduction methods that choose samples heavily depending on repeated training of SVMs. Experiments on both synthetic and real world datasets show that the SR-DSA is capable of reducing the number of samples as well as the time for SVMs training while maintaining high testing accuracy.

Original languageEnglish
Pages (from-to)2772-2781
Number of pages10
JournalNeurocomputing
Volume71
Issue number13-15
DOIs
StatePublished - Aug 2008
Externally publishedYes

Keywords

  • Hierarchical clustering
  • Mahalanobis distance
  • Sample reduction
  • Support vector machines

Fingerprint

Dive into the research topics of 'Selecting valuable training samples for SVMs via data structure analysis'. Together they form a unique fingerprint.

Cite this