1Associate Prof, Dept. of Computer Science and Engg, Jadavpur University, 188 Raja S.C. Mallik Road, Jadavpur, Kolkata-700032, India
2IEEE Member & Electrical Engineer from, Jadavpur University
3Professor, Dept. of Computer Science and Engg, Jadavpur University, 188 Raja S.C. Mallik Road, Jadavpur, Kolkata-700032, India
Online published on 12 May, 2017.
The present work aims to serve the research community by utilizing Case-based Reasoning (CBR) concepts in Data Mining applications to achieve the following: (1) Mining text documents to collect keywords and storing them in a FP-tree structure to generate frequent key-phrases. (2) Using Open Hashing technique to produce an index for the document from associated keywords, and storing in a variation of Separate Chained data structure to resolve collisions. And (3) checking key-words and key-phrases from temporary indexed store of test documents against key-words and key-phrases from relevant documents to test for plagiarism level. For this purpose, key-phrases, which are assumed to be frequent patterns derived as bi-grams and tri-grams of unique key-words, are mined from the documents using frequent pattern generation algorithms such as FP-tree-growth algorithm and Apriori based algorithms. The superiority of FP-tree-growth is experimentally verified and its affinity with the CBR-strategy well established [4]. The threshold experience level of the CBR-system is preliminarily set by including electronic text-books on relevant topics within the document corpus. To facilitate subject-wise indexing, keywords of the test document are mapped onto an index-base to detect measures which guide the classifier module for both retaining a new case as well as retrieving existing ones from the corpus.
Case-Based Reasoning (CBR), Bi-grams, Tri-grams, Frequent Patterns, FP-tree-growth Algorithm, Separate Chaining
(216.73.216.189)