Department of Information Science, University of Madras, Chennai 600 005, E-Mail: g_velumani@yahoo.com; ksragav@hotmail.com
This paper describes an experiment for the automatic identification and extraction of keywords from web resources, specifically HTML document. A natural language parser capable of extracting keywords from HTML texts in a domain would be useful in analyzing document content to support and facilitate information retrieval. A computer program - AutoIndex -was developed to process and extract keywords from HTML texts. Evaluation shows that AutoIndex works fairly well in terms of identification, recall and accuracy. The processing speed of the software is also at acceptable level.