引用本文:田树聪,谢愈,张远龙,周正春,高阳.面向参数化动作空间的多智能体中心化策略梯度分解及其应用.软件学报,2025,36(2):590-607
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览 2407次   下载 3475 本文二维码信息
码上扫一扫!
分享到: 微信 更多
面向参数化动作空间的多智能体中心化策略梯度分解及其应用
田树聪1, 谢愈2, 张远龙2, 周正春1, 高阳3
1.西南交通大学 信息科学与技术学院, 四川 成都 611756;2.国防科技大学 智能科学学院, 湖南 长沙 410073;3.计算机软件新技术国家重点实验室(南京大学), 江苏 南京 210023
摘要:
近年来, 多智能体强化学习方法凭借AlphaStar、AlphaDogFight、AlphaMosaic等成功案例展示出卓越的决策能力以及广泛的应用前景. 在真实环境的多智能体决策系统中, 其任务的决策空间往往是同时具有离散型动作变量和连续型动作变量的参数化动作空间. 这类动作空间的复杂性结构使得传统单一针对离散型或连续型的多智能体强化学习算法不在适用, 因此研究能用于参数化动作空间的多智能体强化学习算法具有重要的现实意义. 提出一种面向参数化动作空间的多智能体中心化策略梯度分解算法, 利用中心化策略梯度分解算法保证多智能体的有效协同, 结合参数化深度确定性策略梯度算法中双头策略输出实现对参数化动作空间的有效耦合. 通过在Hybrid Predator-Prey场景中不同参数设置下的实验结果表明该算法在经典的多智能体参数化动作空间协作任务上具有良好的性能. 此外, 在多巡航导弹协同突防场景中进行算法效能验证, 实验结果表明该算法在多巡航导弹突防这类具有高动态、行为复杂化的协同任务中有效性和可行性.
关键词:  参数化动作空间  多智能体强化学习  中心化策略梯度分解  多巡航导弹突防
DOI:10.13328/j.cnki.jos.007150
分类号:TP18
基金项目:国家自然科学基金(62173336, 92271108)
Factored Multi-agent Centralised Policy Gradient with Parameterized Action Space and Its Application
TIAN Shu-Cong1, XIE Yu2, ZHANG Yuan-Long2, ZHOU Zheng-Chun1, GAO Yang3
1.School of Information Science and Technology, Southwest Jiaotong University, Chengdu 611756, China;2.College of Intelligence Science and Technology, National University of Defense Technology, Changsha 410073, China;3.State Key Laboratory for Novel Software Technology (Nanjing University), Nanjing 210023, China
Abstract:
In recent years, multi-agent reinforcement learning methods have demonstrated excellent decision-making capabilities and broad application prospects in successful cases such as AlphaStar, AlphaDogFight, and AlphaMosaic. In the multi-agent decision-making system in a real-world environments, the decision-making space of its task is often a parameterized action space with both discrete and continuous action variables. The complex structure of this type of action space makes traditional multi-agent reinforcement learning algorithms no longer applicable. Therefore, researching for parameterized action spaces holds important significance in real-world application. This study proposes a factored multi-agent centralised policy gradients algorithm for parameterized action space in multi-agent settings. By utilizing the factored centralised policy gradient algorithm, effective coordination among multi-agent is ensured. After that, the output of the dual-headed policy in the parameterized deep deterministic policy gradient algorithm is employed to achieve effective coupling in the parameterized action space. Experimental results under different parameter settings in the hybrid predator-prey scenario show that the algorithm has good performance on classic multi-agent parameterized action space collaboration tasks. Additionally, the algorithm’s effectiveness and feasibility is validated in a multi-cruise-missile collaborative penetration tasks with complex and high dynamic properties.
Key words:  parameterized action space  multi-agent reinforcement learning  factored centralised policy gradient  multi-cruise-missile collaborative penetration

引用本文:
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览次   下载  
分享到: 微信 更多
摘要:
关键词:  
DOI:
分类号:
基金项目:
Abstract:
Key words: