CyberQ Consulting Private Limited, #622, DLF Tower A, Jasola, New Delhi-110025, India. E-mail: barik011@gmail.com
Online published on 30 March, 2016.
The simply the web, is the most dynamic environment. The web has grown steadily in recent years and his content is changing every day. web is recognized as the largest data source in the world. In this paper, we present a Web Mining process able to discover knowledge in a distributed and heterogeneous multi organization environment. The Web Text Mining process is based on flexible architecture and is implemented by four steps able to examine web content and to extract useful hidden information through mining techniques. An important aspect in Web Mining is played by the automation of extraction rules with proper algorithms. Machine Learning techniques have been successfully applied to Web Mining and Information Extraction tasks thanks to the generalization and adaptation capabilities that are a key requirement on general content, heterogeneous web pages In order to keep the recognition speed high enough for real-world applications an additional algorithm is proposed which lets the approach to boost both in speed and quality.
Web mining, machine learning, unstructured data, and intelligent web agent