Skip to main navigation Skip to search Skip to main content

A case based method to predict optimal k value for k-NN algorithm

  • Yang Zhongguo
  • , Li Hongqi
  • , Zhu Liping
  • , Liu Qiang
  • , Sikandar Ali

Research output: Contribution to journalArticlepeer-review

Abstract

The k-nearest-neighbor classifier is a vital algorithm. In practice, the choice of k is decided by the cross-validation method. We propose a new method for neighborhood size selection based on the data set profile. The distribution of a data set and its intrinsic characteristics are the fundamental factors to the choice of k. A local complexity was computed for each example and a complexity profile was constructed by sorting these local complexity values which try to capture inner structure of a data set. After this, a feature vector was built by combing the local complexity profile and some statistic features of a data set. In addition, a history meta-data set was constructed by using the feature vector as attributes and the optimum k value of data set as the label, which was calculated by using ten cross-validation methods. A predict model was trained based on the historic meta-data set and used to predict optimum k value for a new data set. Some exclusive experiments are conducted to verify the proposed method. The results shows that the local complexity features could reflect the inner structure of a data set which could help find the optimum k for k-NN for different domains.
Original languageEnglish
Pages (from-to)55-65
Number of pages11
JournalJournal of Intelligent and Fuzzy Systems
Volume33
Issue number1
Early online date14 Mar 2017
DOIs
Publication statusPublished - 31 Jul 2017
Externally publishedYes

Keywords

  • k-NN classifier
  • data sets
  • local complexity profile
  • optimum k

Fingerprint

Dive into the research topics of 'A case based method to predict optimal k value for k-NN algorithm'. Together they form a unique fingerprint.

Cite this