引用本文:蔡瑞初,林富艺,陈薇,朱海鹏,郝志峰.隐变量因果模型视角下的策略梯度方差优化.软件学报,2026,37(6):2564-2583
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览 577次   下载 526 本文二维码信息
码上扫一扫!
分享到: 微信 更多
隐变量因果模型视角下的策略梯度方差优化
蔡瑞初1, 林富艺1, 陈薇1, 朱海鹏1, 郝志峰1,2
1.广东工业大学 计算机学院, 广东 广州 510006;2.汕头大学 数学与计算机学院, 广东 汕头 515063
摘要:
深度强化学习已在多个领域取得了显著突破, 其中策略梯度算法因适用于处理非线性和高维状态空间的问题而被广泛采用. 然而, 现有策略梯度算法在实际应用中仍面临高方差问题, 这会导致算法收敛速度变慢, 甚至可能陷入次优解. 针对这一挑战, 从隐变量因果模型的视角提出一种策略梯度方差优化方法. 通过引入隐变量刻画未观测随机信息, 构建并学习隐变量因果模型. 基于隐变量因果模型, 提出因果价值函数, 结合长短期记忆网络, 根据时效性区分衡量未观测随机信息对价值估计的影响作用, 提高动作优势函数预估的准确性, 降低策略梯度方差. 实验表明, 与前沿的同类算法相比, 基于隐变量因果模型的方法在多个任务中更具有优越性和稳定性.
关键词:  深度强化学习  策略梯度  方差优化  隐变量因果模型  因果价值函数
DOI:10.13328/j.cnki.jos.007547
分类号:TP183
基金项目:新一代人工智能国家科技重大专项(2021ZD0111500); 国家优秀青年科学基金(62122022); 国家自然科学基金(62206064)
Variance Optimization of Policy Gradients from Latent Variable Causal Model Perspective
CAI Rui-Chu1, LIN Fu-Yi1, CHEN Wei1, ZHU Hai-Peng1, HAO Zhi-Feng1,2
1.School of Computer Science and Technology, Guangdong University of Technology, Guangzhou 510006, China;2.College of Mathematics and Computer Science, Shantou University, Shantou 515063, China
Abstract:
Deep reinforcement learning has achieved significant breakthroughs in various fields, with policy gradient algorithms widely adopted due to their suitability for handling nonlinear and high-dimensional state spaces. However, in practical applications, existing policy gradient algorithms still suffer from high variance, which slows convergence and may cause suboptimal solutions. To tackle this challenge, a variance optimization method for policy gradients is proposed from a latent variable causal model perspective. By introducing latent variables to characterize unobserved random information, a latent variable causal model is constructed and learned. Utilizing this model, a causal value function is proposed and combined with long short-term memory (LSTM) networks to differentiate the temporal impact of unobserved information on value estimation. This approach improves the accuracy of action advantage function estimation and reduces policy gradient variance. Experiments demonstrate that the proposed latent variable causal model outperforms state-of-the-art algorithms across multiple tasks, with better performance and stability.
Key words:  deep reinforcement learning  policy gradient  variance optimization  latent variable causal model  causal value function

引用本文:
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览次   下载  
分享到: 微信 更多
摘要:
关键词:  
DOI:
分类号:
基金项目:
Abstract:
Key words: