Abstract
The k-nearest-neighbor classifier is a vital algorithm. In practice, the choice of k is decided by the cross-validation method. We propose a new method for neighborhood size selection based on the data set profile. The distribution of a data set and its intrinsic characteristics are the fundamental factors to the choice of k. A local complexity was computed for each example and a complexity profile was constructed by sorting these local complexity values which try to capture inner structure of a data set. After this, a feature vector was built by combing the local complexity profile and some statistic features of a data set. In addition, a history meta-data set was constructed by using the feature vector as attributes and the optimum k value of data set as the label, which was calculated by using ten cross-validation methods. A predict model was trained based on the historic meta-data set and used to predict optimum k value for a new data set. Some exclusive experiments are conducted to verify the proposed method. The results shows that the local complexity features could reflect the inner structure of a data set which could help find the optimum k for k-NN for different domains.
| Original language | English |
|---|---|
| Pages (from-to) | 55-65 |
| Number of pages | 11 |
| Journal | Journal of Intelligent and Fuzzy Systems |
| Volume | 33 |
| Issue number | 1 |
| Early online date | 14 Mar 2017 |
| DOIs | |
| Publication status | Published - 31 Jul 2017 |
| Externally published | Yes |
Keywords
- k-NN classifier
- data sets
- local complexity profile
- optimum k
Fingerprint
Dive into the research topics of 'A case based method to predict optimal k value for k-NN algorithm'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver