|
|
An efficient soft clustering algorithm for web page predictionAbstract: Clustering is the process of organizing objects into groups whose members are similar in some way. It can be considered the most important unsupervised learning problem which deals with finding a structure in a collection of unlabeled data. A cluster is therefore a collection of objects which are “similar” between them and are “dissimilar” to the objects belonging to other clusters. Uniform resource locator (URL) is an addressing scheme used by World Wide Web browsers to locate resources on the Internet. The resource can be any type of file stored on a server, such as a Web page, a text file, a graphics file, or an application program. A web page is a document or information resource that is suitable for the World Wide Web and can be accessed through a web browser. When web pages are stored in a common directory of a web server, they become a website. A website will typically contain a group of web pages that are linked together, or have some other coherent method of navigation. These web pages are linked together using URLs. A web page may be static or dynamic that is it may contain static URLs or dynamic URLs. Out of these URLs some may be accessed very frequently and some may not. Thus the number of URLs in a web page is clustered together based on some similarity measure to get a cluster of most probable URLs to be accessed in near future. This paper explains such an algorithm that clusters URLs present in a web page using soft clustering concept.
|