1Sathyabama University.
2Perunthalaivar Kamarajar Institute of Technology.
The field of text mining has got its attraction by the availability of fields like computational Linguistics, machine learning and the rebirth of bio-informatics but the number of working systems or detailed experimental evaluation are very few. The area of Information extraction (IE) offer the benefits of getting the desired information from multiple documents with less time than it is required to collect from search engines. The purpose of the paper is to develop a domain-independent IE model for extracting entities and facts from technical abstracts. This is done with the construction of patterns using tokens trained with machine learning techniques. The appropriateness of the tokens used for the construction of patterns are trained with the help of categorization techniques like Naïve Bayes, Decision tree, Decision stump, C4.8 and Genetic algorithm. A tool called CTSS is developed using the patterns and the performance of the tool is found to have produced higher recall value than google search engine.
Parsing, Inductive logic programming, Wrapper, IE, recall, precision