Page 263 - 《软件学报》2026年第6期
P. 263
2582 软件学报 2026 年第 37 卷第 6 期
[15] Schulman J, Levine S, Abbeel P, Jordan MI, Moritz P. Trust region policy optimization. In: Proc. of the 32nd Int’l Conf. on Machine
Learning. Lille: JMLR.org, 2015. 1889–1897.
[16] Liu YL, Zhang KQ, Başar T, Yin WT. An improved analysis of (variance-reduced) policy gradient and natural policy gradient methods.
In: Proc. of the 34th Int’l Conf. on Neural Information Processing Systems. Vancouver: Curran Associates Inc., 2020. 639.
[17] Guo JM, Zhang R, Zhang XS, Peng SH, Yi Q, Du ZD, Hu X, Guo Q, Chen YJ. Hindsight value function for variance reduction in
stochastic dynamic environment. In: Proc. of the 30th Int’l Joint Conf. on Artificial Intelligence. Montreal: ijcai.org, 2021. 2476–2482.
[doi: 10.24963/ijcai.2021/341]
[18] Palanisamy P. Hands-on Intelligent Agents with OpenAI Gym: A Step-by-step Guide to Develop AI Agents Using Deep Reinforcement
Learning. Birmingham: Packt Publishing Ltd., 2018.
[19] Schulman J, Wolski F, Dhariwal P, Radford A, Klimov O. Proximal policy optimization algorithms. arXiv:1707.06347, 2017.
[20] Lin HH, Ding WH, Liu ZX, Niu YR, Zhu JC, Niu YM, Zhao D. Safety-aware causal representation for trustworthy offline reinforcement
learning in autonomous driving. IEEE Robotics and Automation Letters, 2024, 9(5): 4639–4646. [doi: 10.1109/LRA.2024.3379805]
[21] Chen W, Zhang J, Zhu H, Xu B, Hao Z, Zhang K, Ye J, Cai R. Causal-aware large language models: Enhancing decision-making through
learning, adapting and acting. In: Proc. of the 34th Int’l Joint Conf. on Artificial Intelligence. Montreal: ijcai.org, 2025. 4292–4300. [doi:
10.24963/IJCAI.2025/478]
[22] Wang JW, Du DH, Tian LL, Chen YK, Li YD, Li YY. ERCI: An explainable experience replay approach with causal inference for deep
reinforcement learning. In: Proc. of the 39th AAAI Conf. on Artificial Intelligence. Philadelphia: AAAI Press, 2025. 27671–27679. [doi:
10.1609/aaai.v39i26.34981]
[23] Cao HY, Feng F, Yang TP, Huo J, Gao Y. Causal information prioritization for efficient reinforcement learning. In: Proc. of the 13th Int’l
Conf. on Learning Representations. Singapore: OpenReview.net, 2025.
[24] Zhang A, McAllister RT, Calandra R, Gal Y, Levine S. Learning invariant representations for reinforcement learning without
reconstruction. In: Proc. of the 9th Int’l Conf. on Learning Representations. OpenReview.net, 2021.
[25] Bennett A, Kallus N, Li LH, Mousavi A. Off-policy evaluation in infinite-horizon reinforcement learning with latent confounders. In:
Proc. of the 24th Int’l Conf. on Artificial Intelligence and Statistics. PMLR, 2021. 1999–2007.
[26] Huang BW, Feng F, Lu CC, Magliacane S, Zhang K. AdaRL: What, where, and how to adapt in transfer reinforcement learning. In: Proc.
of the 10th Int’l Conf. on Learning Representations. OpenReview.net, 2022.
[27] Feng F, Huang BW, Zhang K, Magliacane S. Factored adaptation for non-stationary reinforcement learning. In: Proc. of the 36th Int’l
Conf. on Neural Information Processing Systems. New Orleans: Curran Associates Inc., 2022. 2316.
[28] Gu SX, Lillicrap TP, Ghahramani Z, Turner RE, Levine S. Q-Prop: Sample-efficient policy gradient with an off-policy Critic. In: Proc. of
the 5th Int’l Conf. on Learning Representations. Toulon: OpenReview.net, 2017.
[29] Espeholt L, Soyer H, Munos R, Simonyan K, Mnih V, Ward T, Doron Y, Firoiu V, Harley T, Dunning I, Legg S, Kavukcuoglu K.
IMPALA: Scalable distributed deep-RL with importance weighted actor-learner architectures. In: Proc. of the 35th Int’l Conf. on Machine
Learning. Stockholm: PMLR, 2018. 1406–1415.
[30] Gruslys A, Dabney W, Azar MG, Piot B, Bellemare MG, Munos R. The Reactor: A fast and sample-efficient Actor-Critic agent for
reinforcement learning. In: Proc. of the 6th Int’l Conf. on Learning Representations. Vancouver: OpenReview.net, 2018.
[31] Xu P, Gao F, Gu QQ. An improved convergence analysis of stochastic variance-reduced policy gradient. In: Proc. of the 35th Conf. on
Uncertainty in Artificial Intelligence. Tel Aviv: AUAI Press, 2019. 541–551.
[32] Schulman J, Moritz P, Levine S, Jordan MI, Abbeel P. High-dimensional continuous control using generalized advantage estimation.
arXiv:1506.02438, 2015.
[33] Wu C, Rajeswaran A, Duan Y, Kumar V, Bayen AM, Kakade SM, Mordatch I, Abbeel P. Variance reduction for policy gradient with
action-dependent factorized baselines. In: Proc. of the 6th Int’l Conf. on Learning Representations. Vancouver: OpenReview.net, 2018.
[34] Mao HZ, Venkatakrishnan SB, Schwarzkopf M, Alizadeh M. Variance reduction for reinforcement learning in input-driven environments.
In: Proc. of the 7th Int’l Conf. on Learning Representations. New Orleans: OpenReview.net, 2019.
[35] Pawlowski N, Castro DC, Glocker B. Deep structural causal models for tractable counterfactual inference. In: Proc. of the 34th Int’l Conf.
on Neural Information Processing Systems. Vancouver: Curran Associates Inc., 2020. 73.
附中文参考文献
[5] 林谦, 余超, 伍夏威, 董银昭, 徐昕, 张强, 郭宪. 面向机器人系统的虚实迁移强化学习综述. 软件学报, 2024, 35(2): 711–738. http://
www.jos.org.cn/1000-9825/7006.htm [doi: 10.13328/j.cnki.jos.007006]

