1CSE Department, CMR College of Engineering & Technology, JNTU, Hyderabad, India.
2Department of Computer Science & Engineering, University College of Engineering, Osmania University, Hyderabad-500007, AP, India.
3Department of Computer Science and Engineering, JNTU College of Engineering Kukatpally, Hyderabad, India.
In recent times, the concept of Web Crawling has received remarkable significance owing to the drastic development of the World Wide Web. Huge challenges have been posed by the voluminous amounts of web documents swarming the web to the web search engines making their less appropriate to the users. Additional overheads are created for the search engines by the presence of duplicate and near duplicate web documents in abundance, by which their performance and quality is significantly affected. The web crawling research community has extensively recognized the detection of duplicate and near duplicate web pages. Providing the users with pertinent results for their queries in the first page without duplicate and redundant results is a vital requisite. We have presented a novel and efficient approach for the detection of near duplicate web pages in web crawling in this paper. The near duplicate web pages are detected followed by the storage of crawled web pages in to repositories. The keywords are extracted from the crawled pages initially and on the basis of the extracted keywords, the similarity score between the two pages is calculated. The documents are considered as near duplicates if its similarity scores are lesser than a threshold value. Memory for repositories has been reduced and the search engine quality has been improved owing to the detection.
Web Mining, Web Content Mining, Web Crawling, Web pages, Stemming, Common words, Near duplicate pages, Near duplicate detection