Advances in Computational Sciences and Technology
  • Year: 2008
  • Volume: 1
  • Issue: 2

CTSS: A Tool For Efficient Information Extraction

  • Author:
  • A. Christy1, P. Thambidurai2
  • Total Page Count: 11
  • Page Number: 141 to 151

1Sathyabama University.

2Pondicherry Engineering College.

Abstract

In this era of information technology, the world of text is huge and expanding. Information exists in huge volumes in the form of Internet, Journals, Newspapers, Books, Research papers, Product Manuals, etc. The problem of text mining i.e., discovering useful knowledge from unstructured text is attracting increasing attention. Information extraction (IE) software identifies and removes relevant information from texts, pulling information from a variety of sources, and aggregates it to create a single view. Information extraction is always tied to particular corpora and is poor in recall value. Therefore, making the system as a domain-independent one as well as improving the recall value is an important challenge of IE. In this paper, we propose a domain-independent algorithm for information extraction, called SOFTRULEMINING for extracting the aim, methodology and conclusion specified by authors in technical abstracts. The algorithm is implemented by combining trigram model combined with soft matching rules. A tool CTSS, is implemented using SOFTRULEMINING and is tested with technical abstracts of www.computer.org and www.ansinet.org and we have found that the system has improved its precision value and therefore the recall value after the application of our algorithm SOFTRULEMINING against other search engines.

Keywords

Parsing, Trigram model, Soft matching, information extraction, recall, precision