Information Studies
  • Year: 2005
  • Volume: 11
  • Issue: 3

Automatic Extraction of Keywords from Web Resources

  • Author:
  • G. Velumani, K. S. Raghavan
  • Total Page Count: 10
  • Page Number: 185 to 194

Department of Information Science, University of Madras, Chennai 600 005, E-Mail: g_velumani@yahoo.com; ksragav@hotmail.com

Abstract

This paper describes an experiment for the automatic identification and extraction of keywords from web resources, specifically HTML document. A natural language parser capable of extracting keywords from HTML texts in a domain would be useful in analyzing document content to support and facilitate information retrieval. A computer program - AutoIndex -was developed to process and extract keywords from HTML texts. Evaluation shows that AutoIndex works fairly well in terms of identification, recall and accuracy. The processing speed of the software is also at acceptable level.