| 摘要: |
| 考虑到实验数据的大规模性及不完备性等特点,根据集对分析理论,提出一种新超图模型不完备文本系统的聚类算法,即在超图边的权重中引入了集对的同异反联系度和集对的相似联系度并建立了超图模型,最后应用超图分隔法进行聚类.该算法克服了传统聚类算法的缺陷,更有效地降低了文本空间的维数,提高了不完备文本信息系统聚类的精度和速度.最后的实例说明了该算法的可行性和有效性. |
| 关键词: 不完备信息系统 集对分析方法 高维聚类 超图模型 文本聚类 |
| DOI: |
| 分类号: |
| 基金项目:Supported by the National Natural Science Foundation of China under Grant Nos.60873179, 10971186 (国家自然科学基金); the Foundation of Fujian Province Educational Department of China under Grant No.JB08187 (福建省教育厅B类基金项目) |
|
| Clustering Method for Incomplete Text System Based on Set Pair Analysis |
|
LIN Guo-Ping,LI Shao-Zi
|
| Abstract: |
| This paper presents a novel approach for incomplete text system. Which is based on hypergraph model clustering by using the set-pair analysis, and in which the similar, different and anti-contact connectivity of Set-pair and the similarity value of set-pair are used. After hypergraph model set up, a hypergraph partitioning algorithm is used to find clusters. This new method can eliminate disadvantageous factors and decrease the number of dimensions of the incomplete text data and enhance the speed largely and precision of text clustering. The experimental results show that the algorithm is feasibile and efficient. |
| Key words: incomplete system set-pair analysis high-dimensional clustering hypergragh model text clustering |