Asian Journal of Research in Social Sciences and Humanities
  • Year: 2017
  • Volume: 7
  • Issue: 1

Enhancing Clustering Performance through Filter based Data Reduction

*Assistant Professor, Pandian Saraswathi Yadav Engineering College, India

**Professor, Dhaya College of Engineering, India

Online published on 12 January, 2017.

Abstract

Clustering is an unsupervised learning method used to identify inherent grouping of set of unlabeled data. Such set of groups are termed as Clusters. Grouping of datasets in to clusters involves minimizes the interclass similarity and maximizes the intraclass similarity. Using clustering method the Data Reduction becomes a simple process of identifying a relevant subset in a very big datasets. A comparative study is carried out to identify the performance of filters on different datasets. The various clustering algorithms are tested on those datasets, the results are recorded and analyzed. The Filter is a preprocessing tool which performs certain transformations on the input data, and make the data suitable for specific application. In this paper the various clustering methods along with various datasets are tested to identify the usage of various filters performance. The recorded results proves the K-Means is the best clustering algorithm when compared with Expectation Maximization algorithm.

Keywords

Clustering, EM, K-Means, Data reduction, Filter