1PG Scholar, Department of Information Technology, Mepco Schlenk Engineering College, Sivakasi, Tamilnadu, India. Email: rohniromtech@gmail.com
2Assistant professor, Department of Information Technology, Mepco Schlenk Engineering College, Sivakasi, Tamilnadu, India. Email: mblessa@mepcoeng.ac.in
Online published on 6 April, 2016.
Dealing with dimensions is the great challenge, due to “curse of dimensionality”, for effective outlier detection. In a high dimensional data space, it is difficult to detect most related points and most unrelated points. Outlier is the most unrelated points and in the high dimensional data all data points seemed to be a good outlier, which is a great challenge to identify. In this paper, we propose a concept called Reverse Nearest Neighbor (RNN) method. Here, some points which occur most frequently known as hub-points and some points which occur rarely called anti-hub points are identified in a KNN list. In addition to this we show the relationship between outliers and anti-hub points by correlation factor. Unsupervised learning helps to find the clear outliers; throughout this paper we deal with both synthetic data and real data, to detect the clear outliers. Based on analysis, the proposed work shows that RNN-distance based similarity provides higher percentile score to detect outliers when compared to basic knn approach and ABOD method. Based on this distance based method ID3 approach has done to enhance the better outlier detection.
Outlier, High-Dimensional Data, Reverse Nearest Neighbor, Hub, Antihub