User Browsing Behavior-driven Web Crawling

Minghai Liu, Rui Cai, Ming Zhang, and Lei Zhang

Abstract

To optimize the performance of web crawlers, various measures of page importance have been studied to select and order URLs in crawling. Most sophisticated measures (e.g. breadth-first and PageRank) are based on link structure. In this paper, we treat the problem from another perspective and propose to directly measure page importance through mining user interest and behaviors from web browse logs. Unlike most existing approaches which work on single URL, in this paper, both the log mining and the crawl ordering are performed at the granularity of URL pattern. The proposed URL pattern-based crawl orderings are capable to properly predict the importance of newly created (unseen) URLs. Promising experimental results proved the feasibility of our approach.

Details

Publication typeInproceedings
Published inin Proc. of the 20th ACM Conference on Information and Knowledge Management (CIKM 2011)
PublisherAssociation for Computing Machinery, Inc.
> Publications > User Browsing Behavior-driven Web Crawling