引用本文:彭京,唐常杰,元昌安,李川,胡建军.一种基于概念相似度的数据分类方法.软件学报,2007,18(2):311-322
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览 5696次   下载 7779 本文二维码信息
码上扫一扫!
分享到: 微信 更多
一种基于概念相似度的数据分类方法
彭京1,2, 唐常杰1, 元昌安1, 李川1, 胡建军1
1.四川大学,计算机学院,四川,成都,610065;2.成都市公安局,科技处,四川,成都,610017
摘要:
依据数据属性间的相似信息,提出了一种分类方法.该方法将属性矢量化,属性作为m维空间的基本矢量,数据记录作为属性矢量的和.利用属性间先验的概念相似信息,给出了求取任意属性矢量对的相似距离算法,并将数据间相关度计算转换为属性矢量及其相互投影的公式,从而得到任意两条数据的相关度;利用相关度,提出了一种分类算法.用详实的实验证明了该算法的有效性.
关键词:  数据挖掘  概念相似度  相似距离  属性矢量  分类
DOI:
分类号:
基金项目:Supported by the National Natural Science Foundation of China under Grant No.60473071 (国家自然科学基金); the China Postdoctoral Science Foundation under Grant No.20060400002 (中国博士后科学基金); the Major Science and Technology Project of Sichuan Province of China under
A Data Classification Method Based on Concept Similarity
PENG Jing,TANG Chang-Jie,YUAN Chang-An,LI Chuan,HU Jian-Jun
Abstract:
In this paper, a method of classification is proposed based on the similar information of data properties. The new method assumes that data properties are basic vectors of m dimensions, and each of the data is viewed as a sum vector of all the property-vectors. It suggests a novel distance algorithm to get the distance of every pair of the property based on similar information of the basic property vectors. An algorithm of data classification is also presented based on correlation computing formula composed of property vectors and projections of each other. Efficiency of the new method is proved by extensive experiments.
Key words:  data mining  concept similarity  similar distance  property vector  classification