引用本文:杨 宁,唐常杰,王 悦,陈 瑜,郑皎凌.一种基于时态密度的倾斜分布数据流聚类算法.软件学报,2010,21(5):1031-1041
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览 8394次   下载 8727 本文二维码信息
码上扫一扫!
分享到: 微信 更多
一种基于时态密度的倾斜分布数据流聚类算法
杨 宁1, 唐常杰, 王 悦, 陈 瑜, 郑皎凌
四川大学 计算机学院,四川 成都 610065
摘要:
为解决倾斜分布的数据流聚类这一难题,提出了时态密度概念,给出其度量,揭示了其包括可增量计算在 内的一系列数学性质;设计了时态密度树结构,提高了聚类时的存储和检索效率;设计了能够以实时或异步方式捕捉 数据倾斜分布的数据流时态特征的聚类算法TDCA(temporal density based clustering algorithm),其时间复杂度为 O(c×m×lgm).实验结果表明,该算法不仅有较强的功能,而且具有较好的规模可伸缩性.
关键词:  数据流聚类  时态密度  倾斜分布
DOI:
分类号:
基金项目:Supported by the National Natural Science Foundation of China under Grant No.600773169 (国家自然科学基金); the National KeyTechnology R&D Program in the 11th Five-Year Plan of China under Grant No.2006BAI05A01 (国家“十一五”科技支撑计划)
Clustering Algorithm on Data Stream with Skew Distribution Based on Temporal Density
YANG Ning,TANG Chang-Jie,WANG Yue,CHEN Yu,ZHENG Jiao-Ling
Abstract:
To solve the problem of clustering this paper proposes a concept of temporal density, which reveals a set of mathematical properties, especially the incremental computation. A clustering algorithm named TDCA (temporal density based clustering algorithm) with time complexity of O(c×m×lgm) is created with a tree structure implemented for both storage and retrieve efficiency. TDCA is capable of capturing the temporal features of a data stream with skew data distribution either in real time or on demand. The experimental results show that TDCA is functionable and scalable.
Key words:  data stream clustering  temporal density  skew distribution

引用本文:
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览次   下载  
分享到: 微信 更多
摘要:
关键词:  
DOI:
分类号:
基金项目:
Abstract:
Key words: