International Journal of Marketing and Technology
  • Year: 2014
  • Volume: 4
  • Issue: 10

A study on pattern discovery system for web content mining

  • Author:
  • M. Vasavi
  • Total Page Count: 8
  • Page Number: 105 to 112

Online published on 22 January, 2015.

Abstract

There is an explosive growth of information in the World Wide Web thus posing a challenge to Web users to extract essential knowledge from the Web. Web data Extraction is the process of extracting the information that users are interested in, from Semi-structured or unstructured web pages and saving the information as the XML document or relationship model. In this we describe a complete method for mining news from online news sites. This method navigates across these web sites, extracts news reports from them, and analyzes these reports in order to discover interesting news trends. For instance, it applies dynamic schemes for the extraction of news reports, and domain independent statistical strategies for topic identification and trend analysis. As a whole, our method is an application of web mining that attempts to go beyond straightforward news analysis, trying to understand current society interests and to measure the social importance of ongoing events.

Keywords

Web Data Extraction, digital archives, Cloud computing, text mining, newspapers