引用本文:曹玉娟,牛振东,赵堃,彭学平.基于概念和语义网络的近似网页检测算法.软件学报,2011,22(8):1816-1826
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览 6197次   下载 8588 本文二维码信息
码上扫一扫!
分享到: 微信 更多
基于概念和语义网络的近似网页检测算法
曹玉娟1,2, 牛振东1, 赵堃1, 彭学平1
1.北京理工大学 计算机科学技术学院,北京 100081;2.北京航天飞行控制中心,北京 100094
摘要:
在搜索引擎的检索结果页面中,用户经常会得到内容近似的网页.为了提高检索整体性能和用户满意度,提出了一种基于概念和语义网络的近似网页检测算法DWDCS(near-duplicate webpages detection based on concept and semantic network).改进了经典基于小世界理论提取文档关键词的算法.首先对文档概念进行抽取和归并,不但解决了“表达差异”问题,而且有效降低了语义网络的复杂度;从网络结构的几何特征对其进行分析,同时利用网页的语法和结构信息构建特征向量进行文档相似度的计算,由于无须使用语料库,使得算法天生具有领域无关的优点.实验结果表明,与经典的网页去重算法(I-Match)和单纯依赖词汇共现小世界模型的算法相比,DWDCS 具有很好的抵抗噪声的能力,在大规模实验中获得了准确率>90%和召回率>85%的良好测试结果.良好的时空间复杂度及算法性能不依赖于语料库的优点,使其在大规模网页去重实际应用中获得了良好的效果.
关键词:  网页去重算法  小世界网络  近似网页  均方差
DOI:10.3724/SP.J.1001.2011.03890
分类号:
基金项目:国家自然科学基金(60803050, 60705022); 新世纪优秀人才计划(NCET-06-0161)
Near Duplicated Web Pages Detection Based on Concept and Semantic Network
CAO Yu-Juan1,2, NIU Zhen-Dong1, ZHAO Kun1, PENG Xue-Ping1
1.School of Computer Science and Technology, Beijing Institute of Technology, Beijing 100081, China;2.Beijing Aerospace Command Centre, Beijing 100094, China
Abstract:
Reprinting websites and blogs produces a great deal redundant WebPages. To improve search efficiency and user satisfaction, the near-Duplicate WebPages Detection based on Concept and Semantic network (DWDCS) is proposed. In the course of developing a near-duplicate detection system for a multi-billion pages repository, this paper makes two research contributions. First, the key concept is extracted, instead of the keyphrase, to build Small Word Network (SWN). This not only reduces the complexity of the semantic network, but also resolves the “expression difference” problem. Second, this paper considers both syntactic and semantic information to present and compute the documents’ similarities. In a large-scale test, experimental results demonstrate that this approach outperforms that of both I-Match and keyphrase extraction algorithms based on SWN. Many advantages such as linear time and space complexity, without using a corpus, make the algorithm valuable in actual practice.
Key words:  duplicate removal algorithm  small world network  near duplicated Web page  standard deviation

引用本文:
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览次   下载  
分享到: 微信 更多
摘要:
关键词:  
DOI:
分类号:
基金项目:
Abstract:
Key words: