Page 262 - 《软件学报》2026年第6期
P. 262

蔡瑞初 等: 隐变量因果模型视角下的策略梯度方差优化                                                      2581


                    总结参数影响实验, 如图        8  所示可以发现本算法在合理的超参设置范围内, 都能表现出较为稳定的良好性能.
                 更多需要关注的是经典强化部分的超参设置, 如折扣因子这类的超参设置. 在折扣因子的超参设置上, 使用较高值
                 的折扣因子表现会更优秀, 如         1、0.99.
                  6   总 结

                    本文针对动态环境中存在随机信息的情况, 导致现有策略梯度方法存在的高方差问题, 从隐变量因果模型的
                 视角提出了一种策略梯度方差优化方法              PGVO, 该方法通过引入显变量和隐变量分别刻画可观测状态信息和动态
                 环境随机信息, 分析并构建隐变量因果模型, 然后根据该模型提出因果价值函数. 该价值函数估计考虑了环境随机
                 信息带来的干扰, 以其为基线计算动作优势估计值, 可有效降低动态环境随机信息对策略梯度算法表现性能的影
                 响. 上述步骤的求解利用因果编解码框架推断动态环境随机信息, 并结合长短期记忆网络处理变长的动态环境随
                 机信息, 借助因果价值函数结合可观测状态信息进行价值函数估计. 实验结果表明, 本文方法在不同随机种子和不
                 同强度的动态环境随机信息下表现出卓越的性能, 显著提升了策略梯度算法的收敛速度和稳定性. 未来的工作将
                 考虑非稳态等更复杂场景所带来的挑战, 进一步提高策略梯度方法的适用性同时优化方差.

                 References
                  [1]   Wang X, Wang S, Liang XX, Zhao DW, Huang JC, Xu X, Dai B, Miao QG. Deep reinforcement learning: A survey. IEEE Trans. on
                     Neural Networks and Learning Systems, 2024, 35(4): 5064–5078. [doi: 10.1109/TNNLS.2022.3207346]
                  [2]   Perolat J, de Vylder B, Hennes D, et al. Mastering the game of Stratego with model-free multiagent reinforcement learning. Science,
                     2022, 378(6623): 990–996. [doi: 10.1126/science.add4679]
                  [3]   Li  QY,  Peng  ZH,  Feng  L,  Zhang  QH,  Xue  ZH,  Zhou  BL.  MetaDrive:  Composing  diverse  driving  scenarios  for  generalizable
                     reinforcement learning. IEEE Trans. on Pattern Analysis and Machine Intelligence, 2023, 45(3): 3461–3475. [doi: 10.1109/TPAMI.2022.
                     3190471]
                  [4]   Ibarz J, Tan J, Finn C, Kalakrishnan M, Pastor P, Levine S. How to train your robot with deep reinforcement learning: Lessons we have
                     learned. The Int’l Journal of Robotics Research, 2021, 40(4–5): 698–721. [doi: 10.1177/0278364920987859]
                  [5]   Lin Q, Yu C, Wu XW, Dong YZ, Xu X, Zhang Q, Guo X. Survey on sim-to-real transfer reinforcement learning in robot systems. Ruan
                     Jian Xue Bao/Journal of Software, 2024, 35(2): 711–738 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/7006.htm
                     [doi: 10.13328/j.cnki.jos.007006]
                  [6]   Liu  JW,  Gao  F,  Luo  XL.  Survey  of  deep  reinforcement  learning  based  on  value  function  and  policy  gradient.  Chinese  Journal  of
                     Computers, 2019, 42(6): 1406–1438 (in Chinese with English abstract). [doi: 10.11897/SP.J.1016.2019.01406]
                  [7]   Shen WG, Huan X. Bayesian sequential optimal experimental design for nonlinear models using policy gradient reinforcement learning.
                     Computer Methods in Applied Mechanics and Engineering, 2023, 416: 116304. [doi: 10.1016/j.cma.2023.116304]
                  [8]   Wu CC, Ruan JG, Cui HH, Zhang B, Li TY, Zhang KX. The application of machine learning based energy management strategy in multi-
                     mode plug-in hybrid electric vehicle, part I: Twin delayed deep deterministic policy gradient algorithm design for hybrid mode. Energy,
                     2023, 262: 125084. [doi: 10.1016/j.energy.2022.125084]
                  [9]   Zhang JW, Lü S, Zhang ZH, Yu JY, Gong XY. Survey on deep reinforcement learning methods based on sample efficiency optimization.
                     Ruan Jian Xue Bao/Journal of Software, 2022, 33(11): 4217–4238 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/
                     6391.htm [doi: 10.13328/j.cnki.jos.006391]
                 [10]   Liu  H,  Socher  R,  Xiong  CM.  Taming  MAML:  Efficient  unbiased  meta-reinforcement  learning.  In:  Proc.  of  the  36th  Int’l  Conf.  on
                     Machine Learning. Long Beach: PMLR, 2019. 4061–4071.
                 [11]   Munos R, Stepleton T, Harutyunyan A, Bellemare MG. Safe and efficient off-policy reinforcement learning. In: Proc. of the 30th Int’l
                     Conf. on Neural Information Processing Systems. Barcelona: Curran Associates Inc., 2016. 1054–1062.
                 [12]   Papini M, Binaghi D, Canonaco G, Pirotta M, Restelli M. Stochastic variance-reduced policy gradient. In: Proc. of the 35th Int’l Conf. on
                     Machine Learning. Stockholm: PMLR, 2018. 4023–4032.
                 [13]   Nguyen DT, Kumar A, Lau HC. Policy gradient with value function approximation for collective multiagent planning. In: Proc. of the
                     31st Int’l Conf. on Neural Information Processing Systems. Long Beach: Curran Associates Inc., 2017. 4322–4332.
                 [14]   Thomas  PS,  Brunskill  E.  Data-efficient  off-policy  policy  evaluation  for  reinforcement  learning.  In:  Proc.  of  the  33rd  Int’l  Conf.  on
                     Machine Learning. New York: JMLR.org, 2016. 2139–2148.
   257   258   259   260   261   262   263   264   265   266   267