| 摘要: |
| 为解决倾斜分布的数据流聚类这一难题,提出了时态密度概念,给出其度量,揭示了其包括可增量计算在
内的一系列数学性质;设计了时态密度树结构,提高了聚类时的存储和检索效率;设计了能够以实时或异步方式捕捉
数据倾斜分布的数据流时态特征的聚类算法TDCA(temporal density based clustering algorithm),其时间复杂度为
O(c×m×lgm).实验结果表明,该算法不仅有较强的功能,而且具有较好的规模可伸缩性. |
| 关键词: 数据流聚类 时态密度 倾斜分布 |
| DOI: |
| 分类号: |
| 基金项目:Supported by the National Natural Science Foundation of China under Grant No.600773169 (国家自然科学基金); the National KeyTechnology R&D Program in the 11th Five-Year Plan of China under Grant No.2006BAI05A01 (国家“十一五”科技支撑计划) |
|
| Clustering Algorithm on Data Stream with Skew Distribution Based on Temporal Density |
|
YANG Ning,TANG Chang-Jie,WANG Yue,CHEN Yu,ZHENG Jiao-Ling
|
| Abstract: |
| To solve the problem of clustering this paper proposes a concept of temporal density, which reveals a
set of mathematical properties, especially the incremental computation. A clustering algorithm named TDCA
(temporal density based clustering algorithm) with time complexity of O(c×m×lgm) is created with a tree structure implemented for both storage and retrieve efficiency. TDCA is capable of capturing the temporal features of a data stream with skew data distribution either in real time or on demand. The experimental results show that TDCA is functionable and scalable. |
| Key words: data stream clustering temporal density skew distribution |