引用本文:陈俞,赵素云,李雪峰,陈红,李翠平.基于随机抽样的模糊粗糙约简.软件学报,2017,28(11):2825-2835
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览 2825次   下载 4881 本文二维码信息
码上扫一扫!
分享到: 微信 更多
基于随机抽样的模糊粗糙约简
陈俞1,2, 赵素云1,2, 李雪峰3, 陈红1,2, 李翠平1,2
1.中国人民大学 信息学院, 北京 100872;2.数据工程与知识工程教育部重点实验室(中国人民大学), 北京 100872;3.中国人民大学 环境学院, 北京 100872
摘要:
传统的属性约简由于其时间复杂度和空间复杂度过高,几乎无法应用到大规模的数据集中.将随机抽样引入传统的模糊粗糙集中,使得属性约简的效率大幅度提升.首先,在统计下近似的基础上提出一种统计属性约简的定义.这里的约简不是原有意义上的约简,而是保持基于统计下近似定义的统计辨识度不变的属性子集.然后,采用抽样的方法计算统计辨识度的样本估计值,基于此估计值可以对统计属性重要性进行排序,从而可以设计一种快速的适用于大规模数据的序约简算法.由于随机抽样集以及统计近似概念的引入,该算法从时间和空间上均降低了约简的计算复杂度,同时又保持了数据集中信息含量几乎不变.最后,数值实验将基于随机抽样的序约简算法和两种传统的属性约简算法从以下3个方面进行了对比:计算属性约简时间消耗、计算属性约简空间消耗、约简效果.对比实验验证了基于随机抽样的序约简算法在时间与空间上的优势.
关键词:  模糊粗糙集  随机抽样  属性约简  统计粗糙集
DOI:10.13328/j.cnki.jos.005337
分类号:
基金项目:国家重点研发计划(2016YFB1000702);国家重点基础研究发展计划(973)(2014CB340402);国家高技术研究发展计划(863)(2014AA015204);国家自然科学基金(61772536,61772537,61702522,61532021);国家社会科学基金(12&ZD220);中国人民大学科学研究基金(中央高校基本科研业务费专项资金)(15XNLQ06);国家高等学校学科创新引智计划(111)
Fuzzy Rough Reduction Based on Random Sampling
CHEN Yu1,2, ZHAO Su-Yun1,2, LI Xue-Feng3, CHEN Hong1,2, LI Cui-Ping1,2
1.School of Information, Renmin University of China, Beijing 100872, China;2.Key Laboratory of Data Engineering and Knowledge Engineering(Renmin Universityof China), Ministry of Educaion, Beijing 100872, China;3.School of Environment, Renmin University of China, Beijing 100872, China
Abstract:
Traditional attribute reduction is less effective when applying to large-scale datasets because of its high time and space complexity. In this paper, random sampling is introduced into traditional rough reduction. First, statistical discernibility degree and statistical rough reduction are proposed based on statistical rough approximation. Here the statistical rough reduction is not the traditional reduction any more, it is a subset which keeps the statistical discernibility degree almost invariant. By using random sampling to find the estimated value of statistical discernibility degree, all the condition attributes can be sorted. And then the reduction can be done on the sorted attributes by keeping the statistical discernibility degree almost invariant. Finally, numerical experimental comparison demonstrates that the random sampling based rough reduction is effective on both time and space consumption.
Key words:  fuzzy rough set  random sampling  attribute reduction  statistical rough set

引用本文:
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览次   下载  
分享到: 微信 更多
摘要:
关键词:  
DOI:
分类号:
基金项目:
Abstract:
Key words: