引用本文:张昊迪,陈振浩,陈俊扬,周熠,连德富,伍楷舜,林方真.显式知识推理和深度强化学习结合的动态决策.软件学报,2023,34(8):3821-3835
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览 2967次   下载 5540 本文二维码信息
码上扫一扫!
分享到: 微信 更多
显式知识推理和深度强化学习结合的动态决策
张昊迪1, 陈振浩1, 陈俊扬1, 周熠2, 连德富3, 伍楷舜1, 林方真4
1.深圳大学 计算机与软件学院, 广东 深圳 518052;2.上海脑科学与类脑研究中心, 上海 200031;3.中国科学技术大学 计算机科学与技术学院, 安徽 合肥 230026;4.香港科技大学 计算机科学与工程系, 香港 999077
摘要:
近年来, 深度强化学习在序列决策领域被广泛应用并且效果良好, 尤其在具有高维输入、大规模状态空间的应用场景中优势明显. 然而, 深度强化学习相关方法也存在一些局限, 如缺乏可解释性、初期训练低效与冷启动等问题. 针对这些问题, 提出了一种基于显式知识推理和深度强化学习的动态决策框架, 将显式的知识推理与深度强化学习结合. 该框架通过显式知识表示将人类先验知识嵌入智能体训练中, 让智能体在强化学习中获得知识推理结果的干预, 以提高智能体的训练效率, 并增加模型的可解释性. 将显式知识分为两种, 即启发式加速知识与规避式安全知识. 前者在训练初期干预智能体决策, 加快训练速度; 而后者将避免智能体作出灾难性决策, 使其训练过程更为稳定. 实验表明, 该决策框架在不同强化学习算法上、不同应用场景中明显提高了模型训练效率, 并增加了模型的可解释性.
关键词:  知识表示与推理  可解释性  深度强化学习  动态序列决策
DOI:10.13328/j.cnki.jos.006593
分类号:TP18
基金项目:国家自然科学基金(61806132,U2001207,61872248);广东省自然科学基金(2017A030312008);深圳市自然科学基金(ZDSYS20190902092853047,R2020A045);珠江人才计划(2019ZT08X603);广东省普通高校创新团队项目(2019KCXTD005)
Dynamic Decision Making Based on Explicit Knowledge Reasoning and Deep Reinforcement Learning
ZHANG Hao-Di1, CHEN Zhen-Hao1, CHEN Jun-Yang1, ZHOU Yi2, LIAN De-Fu3, WU Kai-Shun1, LIN Fang-Zhen4
1.College of Computer Science and Software Engineering, Shenzhen University, Shenzhen 518052, China;2.Shanghai Center for Brain Science and Brain-inspired Technology, Shanghai 200031, China;3.School of Computer Science and Technology, University of Science and Technology of China, Hefei 230026, China;4.Department of Computer Science and Engineering, Hong Kong University of Science and Technology, Hong Kong 999077, China
Abstract:
In recent years, deep reinforcement learning has been widely used in sequential decisions with positive effects, and it has outstanding advantages in application scenarios with high-dimensional input and large state spaces. However, deep reinforcement learning faces some limitations such as a lack of interpretability, inefficient initial training, and a cold start. To address these issues, this study proposes a dynamic decision framework combing explicit knowledge reasoning with deep reinforcement learning. The framework successfully embeds the priori knowledge in intelligent agent training via explicit knowledge representation and gets the agent intervened by the knowledge reasoning results during the reinforcement learning, so as to improve the training efficiency and the model’s interpretability. The explicit knowledge in this study is categorized into two kinds, namely, heuristic acceleration knowledge and evasive safety knowledge. The heuristic acceleration knowledge intervenes in the decision of the agent in the initial training to speed up the training, while the evasive safety knowledge keeps the agent from making catastrophic decisions to keep the training process stable. The experimental results show that the proposed framework significantly improves the training efficiency and the model’s interpretability under different application scenarios and reinforcement learning algorithms.
Key words:  knowledge representation and reasoning  interpretability  deep reinforcement learning (DRL)  sequential decision making

引用本文:
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览次   下载  
分享到: 微信 更多
摘要:
关键词:  
DOI:
分类号:
基金项目:
Abstract:
Key words: