###
Journal of Software:2015.26(8):2056-2073

MapReduce集群环境下的数据放置策略
荀亚玲,张继福,秦啸
(太原科技大学 计算机科学与技术学院, 山西 太原 030024;Department of Computer Science and Software Engineering, Auburn University, USA)
Data Placement Strategy for MapReduce Cluster Environment
XUN Ya-Ling,ZHANG Ji-Fu,QIN Xiao
(School of Computer Science and Technology, Taiyuan University of Science and Technology, Taiyuan 030024, China;Department of Computer Science and Software Engineering, Auburn University, USA))
Abstract
Chart / table
Reference
Similar Articles
Article :Browse 2272   Download 2583
Received:April 08, 2014    Revised:December 21, 2014
> 中文摘要: MapReduce是一种适用于大规模数据密集型应用的有效编程模型,具有编程简单、易于扩展、容错性好等特点,已在并行和分布式计算领域得到了广泛且成功的应用.由于MapReduce将计算扩展到大规模的机器集群上,处理数据的合理放置成为影响MapReduce集群系统性能(包括能耗、资源利用率、通信和I/O代价、响应时间、系统的可靠性和吞吐率等)的关键因素之一.首先,对MapReduce编程模型的典型实现——Hadoop缺省的数据放置策略进行分析,并进一步讨论了MapReduce框架下,设计数据放置策略时需考虑的关键问题和衡量数据放置策略的标准;其次,对目前MapReduce集群环境下的数据放置策略优化方法的研究与进展进行了综述和分析;最后,分析和归纳了MapReduce集群环境下数据放置策略的下一步研究工作.
中文关键词: 数据放置  MapReduce  编程模型  能耗  负载均衡
Abstract:As an effective programming model for large-scale data-intensive applications, MapReduce has been widely and successfully applied in the field of parallel and distributed computing, and has the characteristics of good fault-tolerance and easy to implement and extend. Because MapReduce extends computing to the nodes of large-scale cluster system, reasonable placement of processing data has become one of the key factors affecting the performance of MapReduce cluster system, including energy efficiency, resource utilization, communications and I/O throughput, response time, and reliability. This study first analyzes characteristics of the default data placement strategy of Hadoop, which is a typical implementation of MapReduce programming model. Next, it investigates popular data placement strategies for MapReduce cluster computing environments. Finally, it presents future research directions in the area of data placement strategies for MapReduce-based cluster computing systems.
文章编号:     中图分类号:    文献标志码:
基金项目:国家自然科学基金(61272263); NSF CAREER Award(CCF-0845257) 国家自然科学基金(61272263); NSF CAREER Award(CCF-0845257)
Foundation items:
Reference text:

荀亚玲,张继福,秦啸.MapReduce集群环境下的数据放置策略.软件学报,2015,26(8):2056-2073

XUN Ya-Ling,ZHANG Ji-Fu,QIN Xiao.Data Placement Strategy for MapReduce Cluster Environment.Journal of Software,2015,26(8):2056-2073