Page 261 - 《软件学报》2026年第6期
P. 261

2580                                                       软件学报  2026  年第  37  卷第  6  期



                                                               10 4
                           −30                                                    PGVO_hdim_4
                                                               10 3               PGVO_hdim_8
                                                                                  PGVO_hdim_16
                                 (496, −39.99)
                          Cumulative reward  −50  (524, −39.96)  PGVO_hdim_4  alue net loss  V  10 2 1
                           −40
                                                                                  PGVO_hdim_32
                                 (520, −39.95)
                                 (515, −39.99)
                                                               10
                           −60
                                               PGVO_hdim_8
                                               PGVO_hdim_16    10 0
                           −70                 PGVO_hdim_32
                                                               10 −1
                               0    5 000  10 000  15 000  20 000  0    5 000  10 000  15 000  20 000
                                         Episodes                            Episodes
                              (c) 累计奖励: 还原隐变量的信息维度对比             (d) 价值网络损失: 还原隐变量的信息维度对比
                           −30                                 10 4
                                                                                 PGVO_loss_0.6_0.4
                                                               10 3              PGVO_loss_0.7_0.3
                                                                                 PGVO_loss_0.8_0.2
                           −40
                          Cumulative reward  −50  (496, −39.99)  Value net loss  10 2 1
                                (493, −39.98)
                                                                                 PGVO_loss_0.9_0.1
                                (500, −39.98)
                                (510, −39.99)
                                                               10
                                             PGVO_loss_0.6_0.4
                           −60
                                             PGVO_loss_0.7_0.3
                                             PGVO_loss_0.8_0.2  10 0
                           −70               PGVO_loss_0.9_0.1
                                                               10 −1
                               0    5 000  10 000  15 000  20 000  0    5 000  10 000  15 000  20 000
                                         Episodes                            Episodes
                            (e) 累计奖励: 重构损失项和 KL 损失项权重对比        (f) 价值网络损失: 重构损失项和 KL 损失项权重对比
                                                               10 4
                           −30                                                  PGVO_ ppoclip_0.10
                                                                                PGVO_ ppoclip_0.15
                                                               10 3             PGVO_ ppoclip_0.20
                          Cumulative reward  −50  (381, −39.99)  Value net loss  10 2 1
                           −40
                                (496, −39.99)
                                (308, −39.96)
                                                               10
                           −60
                                            PGVO_ ppoclip_0.10
                                                               10 0
                                            PGVO_ ppoclip_0.15
                           −70              PGVO_ ppoclip_0.20
                                                               10 −1
                               0    5 000  10 000  15 000  20 000  0    5 000  10 000  15 000  20 000
                                         Episodes                            Episodes
                                (g) 累计奖励: 近端策略裁剪系数对比              (h) 价值网络损失: 近端策略裁剪系数对比
                           −30                                 10 4             PGVO_gamma_1.00
                                                                                PGVO_gamma_0.99
                           −40                                 10 3             PGVO_gamma_0.95
                          Cumulative reward  −50  (520, −39.97)  PGVO_gamma_1.00  Value net loss  10 2 1
                                                                                PGVO_gamma_0.90
                                (496, −39.99)
                                (470, −39.97)
                           −60
                                (564, −40.00)
                                                               10
                           −70
                           −80              PGVO_gamma_0.99    10 0
                                            PGVO_gamma_0.95
                                            PGVO_gamma_0.90
                           −90                                 10 −1
                              0     5 000  10 000  15 000  20 000  0    5 000  10 000  15 000  20 000
                                         Episodes                            Episodes
                                   (i) 累计奖励: 折扣因子对比                  (j) 价值网络损失: 折扣因子对比
                                           图 8 Thrower 环境噪声消融的性能曲线         (续)
   256   257   258   259   260   261   262   263   264   265   266