International Journal of Managment, IT and Engineering
  • Year: 2012
  • Volume: 2
  • Issue: 11

A de-duplication tool: Implementation of duplicate detection algorithms

  • Author:
  • J. Shana, T. Venkatachalam
  • Total Page Count: 8
  • Page Number: 59 to 66

*Department of MCA, Coimbatore Institute of Technology, Tamilnadu, India

**Department of Physics, Coimbatore Institute of Technology, Tamilnadu, India

Online published on 30 September, 2013.

Abstract

Data quality is an important issue in any data store especially when it is used for analysis. Poor quality of data would drastically affect the decision making process. Cleansing is an activity performed before data is loaded into the warehouse to enhance the quality and consistency of the data. This paper addresses one of the key data quality problem namely detection of duplicates. A duplicate detection tool is built that implements two algorithms namely Sorted Neighborhood method and Token based cleansing method. The tool was tested using student dataset of 118 records and showed an accuracy of 88.09%. This paper is a primitive analysis of how duplicates can be detected in a dataset.

Keywords

data quality, data cleansing, duplicate detection, data warehouse