引用本文:马友忠,孟小峰.云数据管理索引技术研究.软件学报,2015,26(1):145-166
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览 25286次   下载 13795 本文二维码信息
码上扫一扫!
分享到: 微信 更多
云数据管理索引技术研究
马友忠1,2, 孟小峰1
1.中国人民大学 信息学院, 北京 100872;2.洛阳师范学院 信息技术学院, 河南 洛阳 471022
摘要:
数据的爆炸式增长给传统的关系型数据库带来了巨大的挑战,使其在扩展性、容错性等方面遇到了瓶颈.而云计算技术依靠其高扩展性、高可用性、容错性等特点,成为大规模数据管理的有效方案.然而现有的云数据管理系统也存在不足之处,其只能支持基于主键的快速查询,因缺乏索引、视图等机制,所以不能提供高效的多维查询、join等操作,这限制了云计算在很多方面的应用.主要对云数据管理中的索引技术的相关工作进行了深入调研,并作了对比分析,指出了其各自的优点和不足;对在云计算环境下针对海量物联网数据的多维索引技术研究工作进行了简单介绍;最后指出了在云计算环境下针对大数据索引技术的若干挑战性问题.
关键词:  云数据管理  索引  Hadoop  大数据  多维查询
DOI:10.13328/j.cnki.jos.004688
分类号:
基金项目:国家自然科学基金(61379050, 91224008); 国家高技术研究发展计划(863)(2013AA013204); 高等学校博士学科点专项科研基金(20130004130001); 中国人民大学科学研究基金(11XNL010)
Research on Indexing for Cloud Data Management
MA You-Zhong1,2, MENG Xiao-Feng1
1.School of Information, Renmin University of China, Beijing 100872, China;2.School of Information and Technology, Luoyang Normal University, Luoyang 471022, China
Abstract:
The explosive growth of the digital data brings great challenges to the relational database management systems in addressing issues in areas such as scalability and fault tolerance. The cloud computing techniques have been widely used in many applications and become the standard effective approach to manage large scale data because of their high scalability, high availability and fault tolerance. The existing cloud-based data management systems can't efficiently support complex queries such as multi-dimensional queries and join queries because of lacking of index or view techniques, limiting the application of cloud computing in many respects. This paper conducts an in-depth research on the index techniques for cloud data management to highlight their strengths and weaknesses. This paper also introduces its own preliminary work on the index for massive IOT data in cloud environment. Finally, it points out some challenges in the index techniques for big data in cloud environment.
Key words:  cloud data management  index  Hadoop  big data  multi-dimensional query