网页变化与增量搜集技术

微信服务号

微信订阅号

2025年5月17日 6:12 星期六

首页 > 过刊浏览>2006年第17卷第5期 >1051-1067

网页变化与增量搜集技术
DOI:
                        
                    
CSTR:
                        
                    
作者:
                        孟涛孟涛
北京大学,计算机科学技术系,网络与分布式系统实验室,北京,100871
在期刊界中查找
在百度中查找
在本站中查找
王继民王继民
北京大学,计算机科学技术系,网络与分布式系统实验室,北京,100871
在期刊界中查找
在百度中查找
在本站中查找
闫宏飞闫宏飞
北京大学,计算机科学技术系,网络与分布式系统实验室,北京,100871
在期刊界中查找
在百度中查找
在本站中查找

                    
作者单位:
作者简介:
通讯作者:
中图分类号:
基金项目:Supported by the National Natural Science Foundation of China under Grant Nos.60573166,60435020(国家自然科学基金);the National Research Foundation for the Doctoral Program of Higher Education of China under Grant No.20030001076(国家教育部博士点基金)

Web Evolution and Incremental Crawling

Author:

MENG Tao
MENG Tao

在期刊界中查找
在百度中查找
在本站中查找
WANG Ji-Min
WANG Ji-Min

在期刊界中查找
在百度中查找
在本站中查找
YAN Hong-Fei
YAN Hong-Fei

在期刊界中查找
在百度中查找
在本站中查找

Affiliation:

Fund Project:

摘要

图/表

访问统计

参考文献

相似文献

引证文献

资源附件

文章评论

摘要:

互联网络中信息量的快速增长使得增量搜集技术成为网上信息获取的一种有效手段,它可以避免因重复搜集未曾变化的网页而带来的时间和资源上的浪费.网页变化规律的发现和利用是增量搜集技术的一个关键.它用来预测网页的下次变化时间甚至变化程度;在此基础上,增量搜集系统还需要考虑网页的变化频率、变化程度和重要性,选择一种最优的任务调度算法来决定不同网页的搜集频率和相对搜集次序.针对网页变化和增量搜集技术这一主题,对最近几年的研究成果作总结,并介绍最新的研究进展.首先论述对网页变化规律的建模、模型参数估计和估计效率等问题;然后介绍几个著名的增量搜集系统,着重分析它们的任务调度算法;最后,从理论上分析和总结增量搜集系统的最佳任务调度算法及其一个基于启发式策略的近似解,并预测其将来的研究趋势.该工作对增量搜集系统的设计和Web演化规律的研究具有参考意义.

关键词:网页变化;增量搜集;调度策略;研究进展

Abstract:

With the massive and ever increasing pages in the Web, incremental crawling has become a promising method to achieve on-line information. Its main advantage is the resource economization, which comes from the avoidance of downloading unchanged pages. For the precision of change prediction, the evolution of Web is generally studied to find out how pages change. In sum, incremental crawlers often integrate change frequency, change extent, and document quality for each page to determine its relative order as well as its download frequency. In this paper, the researches on Web evolution and incremental crawling in recent years are summarized: First, the change of page is modeled as a Poisson process, and the solutions are given to estimate its parameters, especially the change frequency, and then experimental results are shown. Second, based on the change of pages, three public large-scale incremental crawling systems are introduced, with emphasis on their scheduling policies and strategies to enhance page qualities. Third, theoretical analysis and exploration are performed to find the optimal scheduling policy, three approaches from different points of views are utilized to achieve this object, and a heuristic approximate solution is supplied for the feasibility in practice. Finally, research trends in this area are predicted, and three main issues are listed.

Key words:Web evolution;incremental crawling;scheduling policy;research development

引用本文

孟涛,王继民,闫宏飞.网页变化与增量搜集技术.软件学报,2006,17(5):1051-1067

复制

文章指标

点击次数:
下载次数:
HTML阅读次数:
引用次数:

历史

收稿日期:2005-10-11
最后修改日期:2006-01-12
录用日期:
在线发布日期:
出版日期:

微信服务号

微信订阅号

引用本文

相关视频

分享

文章指标

历史

文章二维码

微信服务号

微信订阅号

引用本文

相关视频

分享

微信扫一扫：分享

文章指标

历史

文章二维码