Central Institute of Indian Languages, Manasagangotri, Mysore 570 006, E-mail:suman@ciil.stpmy.soft.net
Based on a paper to the National Seminar on Information/Knowledge Organization in the Humanities, 9 –11 August 2005, Bangalore
The present study attempts to build a knowledge management system for Hindi documents. The titles of 1299 Hindi documents were selected from the CIIL database from the areas of Linguistics, Education, Sociology and Religion. The database is based on software Virtua with UNICODE support. Using NLP tools software the extraction of keywords, word frequency, expressiveness in titles, metaphor study of 1299 Hindi titles has been done. The selected keywords have been searched in the Hindi word net and Hindi thesaurus for the comparison of its vocabulary. The primary objective of the study is to prepare a tool for domain specific thesauri in the area of Hindi and its related areas.