| 摘要: |
| WWW上的信息极大丰富,如何从巨量的信息中有效地发现有用的信息,是亟待解决的问题,而Web网页的正确分类正是其中的核心问题.针对超文本结构中的结构特征,提出了用NaiveBayes方法协调分别利用超文本页面中的文本信息和结构信息进行分类的方法.经实验验证,与只用单种方法对超文本进行分类的方法相比,综合分类法有效地提高了分类的正确率. |
| 关键词: 超文本 Web 分类 机器学习 互联网 数据挖掘 信息检索 WWW |
| DOI: |
| 分类号: |
| 基金项目:国家自然科学基金资助项目(69675016) |
|
| Using Naive Bayes to Coordinate the Classification of Web Pages |
|
FAN Yan,ZHENG Cheng,WANG Qingyi,CAI Qing sheng,LIU Jie
|
| Abstract: |
| There is a vast source of information in WWW. How to find the useful information from Internet is an exact issue to be solved. The correct classification of Web pages is the core. Based on the structure characteristics of hypertext, the method of Naive Bayes is adopted in this paper to coordinate the two classifiers that use the text document and hypertext structure. Compared with the two separate classifiers, the combining classifier promotes the correctness of Web pages'classification evidently and steadily. |
| Key words: hypertext Web classification machine learning Internet data mining information retrieval WWW |