###
Journal of Software:2018.29(3):839-852

多维图结构聚类的社交关系挖掘算法
李振军,代强强,李荣华,毛睿,乔少杰
(深圳大学 计算机与软件学院, 广东 深圳 518060;成都信息工程大学 网络空间安全学院, 四川 成都 610225)
Social Relationship Mining Algorithm by Multi-Dimensional Graph Structural Clustering
LI Zhen-Jun,DAI Qiang-Qiang,LI Rong-Hua,MAO Rui,QIAO Shao-Jie
(College of Computer Science & Software Engineering, Shenzhen University, Shenzhen 518060, China;School of Cybersecurity, Chengdu University of Information Technology, Chengdu 610225, China)
Abstract
Chart / table
Reference
Similar Articles
Article :Browse 1688   Download 2840
Received:August 02, 2017    Revised:September 05, 2017
> 中文摘要: 社交关系的数据挖掘一直是大图数据研究领域中的热门问题.图聚类算法如SCAN (structural clustering algorithm for network)虽然可以迅速地从海量图数据中获得关系紧密的社区结构,但这类社区往往只表示了社交对象的聚集,无法反馈对象间的真实社交关系,如家庭成员、同事、同学等.要获取对象间真实的社交关系,需要更多维度地挖掘现实中社交对象间复杂的交互关系.对象间的交互维度很多,例如通话、见面、微信、电子邮件等,而传统SCAN等聚类算法仅能够挖掘单维度的交互数据.在研究社交对象间的多维社交关系图数据与传统图结构聚类算法的基础上,提出了一种有效的子空间聚类算法SCA (subspace cluster algorithm),对多维度下子空间的图结构聚类进行研究,目的是探索如何通过图数据挖掘发现对象间真实的社交关系.SCA算法遵循自底向上的原则,能够发现社交图数据中所有子空间的聚类集.为提升SCA的运行速度,利用其子空间聚类的单调性进行了性能优化,进而提出了剪枝算法SCA+.最后进行了大规模的性能测试实验以及真实数据的案例研究,其结果验证了算法的效率和效用.
中文关键词: 图聚类  多维图数据  社交关系  子空间
Abstract:Social relationship mining is a hot topic in the area of massive graph analysis. Graph clustering algorithms such as SCAN (structural clustering algorithm for networks) can quickly discover the communities from the massive graph data. However, relationships in these communities fail to reflect the ‘real’ social information such as family, colleagues and classmates. In reality, social data is very complex, and there are many types of interaction among each individual, such as calling, meeting, chatting in WeChat, and sending emails. However, traditional SCAN algorithm can only handle single dimensional graph data. Based on the study of multidimensional social graph data and traditional clustering algorithms, this paper first proposes an efficient subspace clustering algorithm named SCA by mining multi-dimensional clusters in subspaces as a mean to explore real social relationships. SCA follows the bottom-up principle and can discover the set of clusters from the social graph data in all dimensions. To improve the efficiency of SCA, the paper also develops a pruning algorithm called SCA+ based on the monotonicity of subspace clustering. Extensive experiments on several real-world multi-dimensional graph data demonstrate the efficiency and effectiveness of the proposed algorithms.
文章编号:     中图分类号:    文献标志码:
基金项目:国家自然科学基金(61402292,61772091);国家自然科学基金广东省联合基金(U1301252);教育部人文社会科学研究规划基金(15YJAZH058) 国家自然科学基金(61402292,61772091);国家自然科学基金广东省联合基金(U1301252);教育部人文社会科学研究规划基金(15YJAZH058)
Foundation items:National Natural Science Foundation of China (61402292, 61772091);National Natural Science Foundation of China Guangdong Joint Fund Project (U1301252);Planning Foundation for Humanities and Social Sciences of Ministry of Education of China (15YJAZH058)
Reference text:

李振军,代强强,李荣华,毛睿,乔少杰.多维图结构聚类的社交关系挖掘算法.软件学报,2018,29(3):839-852

LI Zhen-Jun,DAI Qiang-Qiang,LI Rong-Hua,MAO Rui,QIAO Shao-Jie.Social Relationship Mining Algorithm by Multi-Dimensional Graph Structural Clustering.Journal of Software,2018,29(3):839-852