Page 263 - 《软件学报》2026年第6期
P. 263

2582                                                       软件学报  2026  年第  37  卷第  6  期


                 [15]   Schulman J, Levine S, Abbeel P, Jordan MI, Moritz P. Trust region policy optimization. In: Proc. of the 32nd Int’l Conf. on Machine
                     Learning. Lille: JMLR.org, 2015. 1889–1897.
                 [16]   Liu YL, Zhang KQ, Başar T, Yin WT. An improved analysis of (variance-reduced) policy gradient and natural policy gradient methods.
                     In: Proc. of the 34th Int’l Conf. on Neural Information Processing Systems. Vancouver: Curran Associates Inc., 2020. 639.
                 [17]   Guo JM, Zhang R, Zhang XS, Peng SH, Yi Q, Du ZD, Hu X, Guo Q, Chen YJ. Hindsight value function for variance reduction in
                     stochastic dynamic environment. In: Proc. of the 30th Int’l Joint Conf. on Artificial Intelligence. Montreal: ijcai.org, 2021. 2476–2482.
                     [doi: 10.24963/ijcai.2021/341]
                 [18]   Palanisamy P. Hands-on Intelligent Agents with OpenAI Gym: A Step-by-step Guide to Develop AI Agents Using Deep Reinforcement
                     Learning. Birmingham: Packt Publishing Ltd., 2018.
                 [19]   Schulman J, Wolski F, Dhariwal P, Radford A, Klimov O. Proximal policy optimization algorithms. arXiv:1707.06347, 2017.
                 [20]   Lin HH, Ding WH, Liu ZX, Niu YR, Zhu JC, Niu YM, Zhao D. Safety-aware causal representation for trustworthy offline reinforcement
                     learning in autonomous driving. IEEE Robotics and Automation Letters, 2024, 9(5): 4639–4646. [doi: 10.1109/LRA.2024.3379805]
                 [21]   Chen W, Zhang J, Zhu H, Xu B, Hao Z, Zhang K, Ye J, Cai R. Causal-aware large language models: Enhancing decision-making through
                     learning, adapting and acting. In: Proc. of the 34th Int’l Joint Conf. on Artificial Intelligence. Montreal: ijcai.org, 2025. 4292–4300. [doi:
                     10.24963/IJCAI.2025/478]
                 [22]   Wang JW, Du DH, Tian LL, Chen YK, Li YD, Li YY. ERCI: An explainable experience replay approach with causal inference for deep
                     reinforcement learning. In: Proc. of the 39th AAAI Conf. on Artificial Intelligence. Philadelphia: AAAI Press, 2025. 27671–27679. [doi:
                     10.1609/aaai.v39i26.34981]
                 [23]   Cao HY, Feng F, Yang TP, Huo J, Gao Y. Causal information prioritization for efficient reinforcement learning. In: Proc. of the 13th Int’l
                     Conf. on Learning Representations. Singapore: OpenReview.net, 2025.
                 [24]   Zhang  A,  McAllister  RT,  Calandra  R,  Gal  Y,  Levine  S.  Learning  invariant  representations  for  reinforcement  learning  without
                     reconstruction. In: Proc. of the 9th Int’l Conf. on Learning Representations. OpenReview.net, 2021.
                 [25]   Bennett A, Kallus N, Li LH, Mousavi A. Off-policy evaluation in infinite-horizon reinforcement learning with latent confounders. In:
                     Proc. of the 24th Int’l Conf. on Artificial Intelligence and Statistics. PMLR, 2021. 1999–2007.
                 [26]   Huang BW, Feng F, Lu CC, Magliacane S, Zhang K. AdaRL: What, where, and how to adapt in transfer reinforcement learning. In: Proc.
                     of the 10th Int’l Conf. on Learning Representations. OpenReview.net, 2022.
                 [27]   Feng F, Huang BW, Zhang K, Magliacane S. Factored adaptation for non-stationary reinforcement learning. In: Proc. of the 36th Int’l
                     Conf. on Neural Information Processing Systems. New Orleans: Curran Associates Inc., 2022. 2316.
                 [28]   Gu SX, Lillicrap TP, Ghahramani Z, Turner RE, Levine S. Q-Prop: Sample-efficient policy gradient with an off-policy Critic. In: Proc. of
                     the 5th Int’l Conf. on Learning Representations. Toulon: OpenReview.net, 2017.
                 [29]   Espeholt L, Soyer H, Munos R, Simonyan K, Mnih V, Ward T, Doron Y, Firoiu V, Harley T, Dunning I, Legg S, Kavukcuoglu K.
                     IMPALA: Scalable distributed deep-RL with importance weighted actor-learner architectures. In: Proc. of the 35th Int’l Conf. on Machine
                     Learning. Stockholm: PMLR, 2018. 1406–1415.
                 [30]   Gruslys A, Dabney W, Azar MG, Piot B, Bellemare MG, Munos R. The Reactor: A fast and sample-efficient Actor-Critic agent for
                     reinforcement learning. In: Proc. of the 6th Int’l Conf. on Learning Representations. Vancouver: OpenReview.net, 2018.
                 [31]   Xu P, Gao F, Gu QQ. An improved convergence analysis of stochastic variance-reduced policy gradient. In: Proc. of the 35th Conf. on
                     Uncertainty in Artificial Intelligence. Tel Aviv: AUAI Press, 2019. 541–551.
                 [32]   Schulman J, Moritz P, Levine S, Jordan MI, Abbeel P. High-dimensional continuous control using generalized advantage estimation.
                     arXiv:1506.02438, 2015.
                 [33]   Wu C, Rajeswaran A, Duan Y, Kumar V, Bayen AM, Kakade SM, Mordatch I, Abbeel P. Variance reduction for policy gradient with
                     action-dependent factorized baselines. In: Proc. of the 6th Int’l Conf. on Learning Representations. Vancouver: OpenReview.net, 2018.
                 [34]   Mao HZ, Venkatakrishnan SB, Schwarzkopf M, Alizadeh M. Variance reduction for reinforcement learning in input-driven environments.
                     In: Proc. of the 7th Int’l Conf. on Learning Representations. New Orleans: OpenReview.net, 2019.
                 [35]   Pawlowski N, Castro DC, Glocker B. Deep structural causal models for tractable counterfactual inference. In: Proc. of the 34th Int’l Conf.
                     on Neural Information Processing Systems. Vancouver: Curran Associates Inc., 2020. 73.

                 附中文参考文献
                 [5]   林谦, 余超, 伍夏威, 董银昭, 徐昕, 张强, 郭宪. 面向机器人系统的虚实迁移强化学习综述. 软件学报, 2024, 35(2): 711–738. http://
                    www.jos.org.cn/1000-9825/7006.htm [doi: 10.13328/j.cnki.jos.007006]
   258   259   260   261   262   263   264   265   266   267   268