引用本文:陈奕宇,霍静,丁天雨,高阳.元强化学习研究综述.软件学报,2024,35(4):1618-1650
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览 6727次   下载 13147 本文二维码信息
码上扫一扫!
分享到: 微信 更多
元强化学习研究综述
陈奕宇1,2, 霍静1,2, 丁天雨3, 高阳1,2
1.南京大学 计算机科学与技术系, 江苏 南京 210043;2.计算机软件新技术国家重点实验室(南京大学), 江苏 南京 210043;3.Applied Sciences Group, Microsoft, Redmond, WA 98034, USA
摘要:
近年来, 深度强化学习(deep reinforcement learning, DRL)已经在诸多序贯决策任务中取得瞩目成功, 但当前, 深度强化学习的成功很大程度依赖于海量的学习数据与计算资源, 低劣的样本效率和策略通用性是制约其进一步发展的关键因素. 元强化学习(meta-reinforcement learning, Meta-RL)致力于以更小的样本量适应更广泛的任务, 其研究有望缓解上述限制从而推进强化学习领域发展. 以元强化学习工作的研究对象与适用场景为脉络, 对元强化学习领域的研究进展进行了全面梳理: 首先, 对深度强化学习、元学习背景做基本介绍; 然后, 对元强化学习作形式化定义及常见的场景设置总结, 并从元强化学习研究成果的适用范围角度展开介绍元强化学习的现有研究进展; 最后, 分析了元强化学习领域的研究挑战与发展前景.
关键词:  元强化学习  强化学习  深度强化学习  元学习
DOI:10.13328/j.cnki.jos.007011
分类号:
基金项目:科技创新2030—“新一代人工智能”重大项目(2021ZD0113303);国家自然科学基金(62192783, 62276128)
Survey of Meta-reinforcement Learning Research
CHEN Yi-Yu1,2, HUO Jing1,2, DING Tian-Yu3, GAO Yang1,2
1.Department of Computer Science and Technology, Nanjing University, Nanjing 210043, China;2.State Key Laboratory for Novel Software Technology(Nanjing University), Nanjing 210043, China;3.Applied Sciences Group, Microsoft, Redmond, WA 98034, USA
Abstract:
In recent years, deep reinforcement learning (DRL) has achieved remarkable success in many sequential decision-making tasks. However, the current success of deep reinforcement learning heavily relies on massive learning data and computing resources. The poor sample efficiency and strategy generalization ability are the key factors restricting DRL’s further development. Meta-reinforcement learning (Meta-RL) studies to adapt to a wider range of tasks with a smaller sample size. Related researches are expected to alleviate the above limitations and promote the development of reinforcement learning. Taking the scope of research object and application range of current research works, this study comprehensively combs the research progress in the field of meta-reinforcement learning. Firstly, a basic introduction is given to deep reinforcement learning and the background of meta-reinforcement learning. Then, meta-reinforcement learning is formally defined and common scene settings are summarized, and the current research progress of meta-reinforcement learning is also introduced from the perspective of application range of the research results. Finally, the research challenges and potential future development directions are discussed.
Key words:  meta-reinforcement learning  reinforcement learning  deep reinforcement learning  meta-learning

引用本文:
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览次   下载  
分享到: 微信 更多
摘要:
关键词:  
DOI:
分类号:
基金项目:
Abstract:
Key words: