Page 262 - 《软件学报》2026年第6期
P. 262
蔡瑞初 等: 隐变量因果模型视角下的策略梯度方差优化 2581
总结参数影响实验, 如图 8 所示可以发现本算法在合理的超参设置范围内, 都能表现出较为稳定的良好性能.
更多需要关注的是经典强化部分的超参设置, 如折扣因子这类的超参设置. 在折扣因子的超参设置上, 使用较高值
的折扣因子表现会更优秀, 如 1、0.99.
6 总 结
本文针对动态环境中存在随机信息的情况, 导致现有策略梯度方法存在的高方差问题, 从隐变量因果模型的
视角提出了一种策略梯度方差优化方法 PGVO, 该方法通过引入显变量和隐变量分别刻画可观测状态信息和动态
环境随机信息, 分析并构建隐变量因果模型, 然后根据该模型提出因果价值函数. 该价值函数估计考虑了环境随机
信息带来的干扰, 以其为基线计算动作优势估计值, 可有效降低动态环境随机信息对策略梯度算法表现性能的影
响. 上述步骤的求解利用因果编解码框架推断动态环境随机信息, 并结合长短期记忆网络处理变长的动态环境随
机信息, 借助因果价值函数结合可观测状态信息进行价值函数估计. 实验结果表明, 本文方法在不同随机种子和不
同强度的动态环境随机信息下表现出卓越的性能, 显著提升了策略梯度算法的收敛速度和稳定性. 未来的工作将
考虑非稳态等更复杂场景所带来的挑战, 进一步提高策略梯度方法的适用性同时优化方差.
References
[1] Wang X, Wang S, Liang XX, Zhao DW, Huang JC, Xu X, Dai B, Miao QG. Deep reinforcement learning: A survey. IEEE Trans. on
Neural Networks and Learning Systems, 2024, 35(4): 5064–5078. [doi: 10.1109/TNNLS.2022.3207346]
[2] Perolat J, de Vylder B, Hennes D, et al. Mastering the game of Stratego with model-free multiagent reinforcement learning. Science,
2022, 378(6623): 990–996. [doi: 10.1126/science.add4679]
[3] Li QY, Peng ZH, Feng L, Zhang QH, Xue ZH, Zhou BL. MetaDrive: Composing diverse driving scenarios for generalizable
reinforcement learning. IEEE Trans. on Pattern Analysis and Machine Intelligence, 2023, 45(3): 3461–3475. [doi: 10.1109/TPAMI.2022.
3190471]
[4] Ibarz J, Tan J, Finn C, Kalakrishnan M, Pastor P, Levine S. How to train your robot with deep reinforcement learning: Lessons we have
learned. The Int’l Journal of Robotics Research, 2021, 40(4–5): 698–721. [doi: 10.1177/0278364920987859]
[5] Lin Q, Yu C, Wu XW, Dong YZ, Xu X, Zhang Q, Guo X. Survey on sim-to-real transfer reinforcement learning in robot systems. Ruan
Jian Xue Bao/Journal of Software, 2024, 35(2): 711–738 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/7006.htm
[doi: 10.13328/j.cnki.jos.007006]
[6] Liu JW, Gao F, Luo XL. Survey of deep reinforcement learning based on value function and policy gradient. Chinese Journal of
Computers, 2019, 42(6): 1406–1438 (in Chinese with English abstract). [doi: 10.11897/SP.J.1016.2019.01406]
[7] Shen WG, Huan X. Bayesian sequential optimal experimental design for nonlinear models using policy gradient reinforcement learning.
Computer Methods in Applied Mechanics and Engineering, 2023, 416: 116304. [doi: 10.1016/j.cma.2023.116304]
[8] Wu CC, Ruan JG, Cui HH, Zhang B, Li TY, Zhang KX. The application of machine learning based energy management strategy in multi-
mode plug-in hybrid electric vehicle, part I: Twin delayed deep deterministic policy gradient algorithm design for hybrid mode. Energy,
2023, 262: 125084. [doi: 10.1016/j.energy.2022.125084]
[9] Zhang JW, Lü S, Zhang ZH, Yu JY, Gong XY. Survey on deep reinforcement learning methods based on sample efficiency optimization.
Ruan Jian Xue Bao/Journal of Software, 2022, 33(11): 4217–4238 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/
6391.htm [doi: 10.13328/j.cnki.jos.006391]
[10] Liu H, Socher R, Xiong CM. Taming MAML: Efficient unbiased meta-reinforcement learning. In: Proc. of the 36th Int’l Conf. on
Machine Learning. Long Beach: PMLR, 2019. 4061–4071.
[11] Munos R, Stepleton T, Harutyunyan A, Bellemare MG. Safe and efficient off-policy reinforcement learning. In: Proc. of the 30th Int’l
Conf. on Neural Information Processing Systems. Barcelona: Curran Associates Inc., 2016. 1054–1062.
[12] Papini M, Binaghi D, Canonaco G, Pirotta M, Restelli M. Stochastic variance-reduced policy gradient. In: Proc. of the 35th Int’l Conf. on
Machine Learning. Stockholm: PMLR, 2018. 4023–4032.
[13] Nguyen DT, Kumar A, Lau HC. Policy gradient with value function approximation for collective multiagent planning. In: Proc. of the
31st Int’l Conf. on Neural Information Processing Systems. Long Beach: Curran Associates Inc., 2017. 4322–4332.
[14] Thomas PS, Brunskill E. Data-efficient off-policy policy evaluation for reinforcement learning. In: Proc. of the 33rd Int’l Conf. on
Machine Learning. New York: JMLR.org, 2016. 2139–2148.

