Department of Computer Science and Engineering, Anna University, Chennai, India
Online published on 1 June, 2016.
In this paper, we propose new information retrieval model with temporal nature for effective information extraction from social networks. For this purpose, we have collected tweets from five thousand users for a period of one month and hotel information related web documents. We considered the hotel related words such as charge, room, television, air condition, food, service, good, bad, excellent, transportation, nearest railway station, bus stand, airport and temples for identifying the user opinion. Based on synonym analysis, we select features for positive sentiments and negative sentiments by proposing a new feature selection algorithm using keyword frequency and semantic analysis. In this paper, we propose a new Latent Dirichlet Allocation (LDA) based Information Retrieval (TLDAIR) model for effective information retrieval in social networks. The proposed model uses clustering technique for grouping the contents of web documents or short messages and rank them according to the relevancy of the given word and also use the Latent Dirichlet Allocation (LDA) for identifying the suitable word for the particular time period which are matched semantically with web document contents in temporal nature. The main advantage of the proposed model is to retrieve the user expected information effectively. From the experiments conducted in this work, it is observed that the groups formed by relevance of documents provided better accuracy than the existing models.
Latent Dirichlet Allocation, Information retrieval, Query, Relevancy, Social Networks