引用本文:黄文振,尹奇跃,张俊格,黄凯奇.基于模型的强化学习中可学习的样本加权机制.软件学报,2023,34(6):2765-2775
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览 1599次   下载 4177 本文二维码信息
码上扫一扫!
分享到: 微信 更多
基于模型的强化学习中可学习的样本加权机制
黄文振1,2, 尹奇跃1,2, 张俊格1,2, 黄凯奇1,2,3
1.中国科学院大学 人工智能学院, 北京 100049;2.中国科学院 自动化研究所 智能系统与工程研究中心, 北京 100190;3.中国科学院 脑科学与智能技术卓越创新中心, 上海 200031
摘要:
基于模型的强化学习方法利用已收集的样本对环境进行建模并使用构建的环境模型生成虚拟样本以辅助训练,因而有望提高样本效率.但由于训练样本不足等问题,构建的环境模型往往是不精确的,其生成的样本也会因携带的预测误差而对训练过程产生干扰.针对这一问题,提出了一种可学习的样本加权机制,通过对生成样本重加权以减少它们对训练过程的负面影响.该影响的量化方法为,先使用待评估样本更新价值和策略网络,再在真实样本上计算更新前后的损失值,使用损失值的变化量来衡量待评估样本对训练过程的影响.实验结果表明,按照该加权机制设计的强化学习算法在多个任务上均优于现有的基于模型和无模型的算法.
关键词:  基于模型的强化学习  模型误差  元学习  强化学习  深度学习
DOI:10.13328/j.cnki.jos.006489
分类号:TP181
基金项目:国家自然科学基金(61876181,61673375);北京市科技创新计划(Z19110000119043);中国科学院青年创新促进会项目;中国科学院项目(QYZDB-SSW-JSC006)
Learnable Weighting Mechanism in Model-based Reinforcement Learning
HUANG Wen-Zhen1,2, YIN Qi-Yue1,2, ZHANG Jun-Ge1,2, HUANG Kai-Qi1,2,3
1.School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China;2.Center for Research on Intelligent System and Engineering (CRISE), Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China;3.Center for Excellence in Brain Science and Intelligence Technology, Chinese Academy of Sciences, Shanghai 200031, China
Abstract:
Model-based reinforcement learning methods train a model to simulate the environment by using the collected samples and utilize the imaginary samples generated by the model to optimize the policy, thus they have potential to improve sample efficiency. Nevertheless, due to the shortage of training samples, the environment model is often inaccurate, and the imaginary samples generated by it would be deleterious for the training process. For this reason, a learnable weighting mechanism is proposed which can reduce the negative effect on the training process by weighting the generated samples. The effect of the imaginary samples on the training process is quantified through calculating the difference between the losses on the real samples before and after updating value and policy networks by the imaginary samples. The experimental results show that the reinforcement learning algorithm using the weighting mechanism is superior to existing model-based and model-free algorithms in multiple tasks.
Key words:  model-based reinforcement learning  model-bias  meta-learning  reinforcement learning  deep learning

引用本文:
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览次   下载  
分享到: 微信 更多
摘要:
关键词:  
DOI:
分类号:
基金项目:
Abstract:
Key words: