Sri Ramakrishna Engineering College, Anna University, Coimbatore-22, Tamilnadu, India.
The abundant amount of data produced and the requirement to merge data from more than one source had resulted in a challenging issue of the efficient detection of duplicate records in databases. Entities possess two or more denotations in real world databases. Generally, duplicate records comprise of errors and are devoid of a common shared key thereby making the task of duplicate matching tedious. A wide variety of methodologies for the identification of duplicate records were projected by numerous researchers. A comprehensive review of the duplicate record detection techniques from significant research works is presented in this paper. An extensive review of the existing literature in duplicate detection of general records in large databases is presented along with the classification. Additionally, a brief introduction about duplicate records detection is presented as well.
Databases, Data cleaning, Data Integration, Duplicate Records, Duplicate detection, Similarity metrics, Duplicate Record Detection Tools, Fuzzy Approach