Page 49 - 《软件学报》2026年第4期
P. 49

1490                                                       软件学报  2026  年第  37  卷第  4  期


                  4.5   小 结

                    通过上述一系列实验分析, 我们观察到, 对抗训练过程中鲁棒性与干净准确率之间存在显著的动态竞争关系.
                 这一现象在模型训练的特定阶段            (通常发生在中后期) 尤为突出, 表现为鲁棒性的提升伴随干净准确率的下降. 这
                 一现象说明, 在模型容量有限的情况下, 其对于对抗样本的鲁棒性与干净样本的泛化能力也是有限的, 两者不能够
                 同时持续增长, 体现了权衡学习的重要性.
                  5   总 结

                    本文提出了一种对抗训练方式           TRG-ASO, 借助遗传算法和攻击策略等概念, 通过对双层优化问题的内层进行
                 改进, 在对抗训练的不同阶段自适应地获取最适合当前模型的攻击策略, 使得模型的鲁棒性和泛化性取得了一个
                 良好的平衡. 相较于标准对抗训练, TRG-ASO          方法训练得到的模型鲁棒性更强, 收敛速度更快; 同时, 攻击策略的
                 变化情况为模型鲁棒性和泛化性的变化趋势做出了合理解释; 此外, 通过记录适应度函数值的变化, 可以分析模型
                 的收敛程度和收敛效果, 从而为模型的早停作出引导, 避免了其对训练数据的过度拟合. 大量实验证明, TRG-ASO
                 具有自适应性, 同时可增强模型的可信性.

                 References
                  [1]   He KM, Zhang XY, Ren SQ, Sun J. Deep residual learning for image recognition. In: Proc. of the 2016 IEEE Conf. on Computer Vision
                     and Pattern Recognition. Las Vegas: IEEE, 2016. 770–778. [doi: 10.1109/CVPR.2016.90]
                  [2]   Collobert R, Weston J. A unified architecture for natural language processing: Deep neural networks with multitask learning. In: Proc. of
                     the 25th Int’l Conf. on Machine Learning. Helsinki: ACM, 2008. 160–167. [doi: 10.1145/1390156.1390177]
                  [3]   Mikolov T, Chen K, Corrado G, Dean J. Efficient estimation of word representations in vector space. arXiv:1301.3781, 2013.
                  [4]   Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I. Attention is all you need. In: Proc. of the
                     31st Conf. on Neural Information Processing Systems. Long Beach: Curran Associates Inc., 2017. 6000–6010. [doi: 10.5555/3295222.
                     3295349]
                  [5]   Szegedy C, Zaremba W, Sutskever I, Bruna J, Erhan D, Goodfellow I, Fergus R. Intriguing properties of neural networks. In: Proc. of the
                     2nd Int’l Conf. on Learning Representations. Banff, 2014. 14–16. [doi: 10.48550/arXiv.1312.6199]
                  [6]   Madry A, Makelov A, Schmidt L, Tsipras D, Vladu A. Towards deep learning models resistant to adversarial attacks. arXiv:1706.06083,
                     2017.
                  [7]   Zhang HY, Yu YD, Jiao JT, Xing E, El Ghaoui L, Jordan M. Theoretically principled trade-off between robustness and accuracy. In:
                     Proc. of the 36th Int’l Conf. on Machine Learning. Long Beach: PMLR, 2019. 7472–7482. [doi: 10.48550/arXiv.1901.08573]
                  [8]   Wang  YS,  Zou  DF,  Yi  JF,  Bailey  J,  Ma  XJ,  Gu  QQ.  Improving  adversarial  robustness  requires  revisiting  misclassified  examples.
                     arXiv:2103.08307, 2021.
                  [9]   Wu HM, Shi WL, Zhang CK, Gu B. Self-adaptive perturbation radii for adversarial training. In: Proc. of the 29th ACM SIGKDD Conf.
                     on Knowledge Discovery and Data Mining. Long Beach: ACM, 2023. 2570–2581. [doi: 10.1145/3580305.3599495]
                 [10]   Yang S, Xu C. One size does not fit all: Data-adaptive adversarial training. In: Proc. of the 17th European Conf. on Computer Vision. Tel
                     Aviv: Springer, 2022. 70–85. [doi: 10.1007/978-3-031-20065-6_5]
                 [11]   Yu CJ, Han B, Gong MM, Shen L, Ge SM, Du B, Liu TL. Robust weight perturbation for adversarial training. In: Proc. of the 31st Int’l
                     Joint Conf. on Artificial Intelligence. Vienna, 2022. 3688–3694. [doi: 10.24963/ijcai.2022/512]
                 [12]   Wang WX, Wang CL, Qi HH, Ye MH, Qian XL, Wang P, Zhang YN. Sustainable self-evolution adversarial training. In: Proc. of the
                     32nd ACM Int’l Conf. on Multimedia. Melbourne: ACM, 2024. 9799–9808. [doi: 10.1145/3664647.3681077]
                 [13]   Goodfellow IJ, Shlens J, Szegedy C. Explaining and harnessing adversarial examples. arXiv:1412.6572, 2014.
                 [14]   Kurakin A, Goodfellow IJ, Bengio S. Adversarial examples in the physical world. arXiv:1607.02533, 2016.
                 [15]   Dong YP, Liao FZ, Pang TY, Su H, Zhu J, Hu XL, Li JG. Boosting adversarial attacks with momentum. In: Proc. of the 2018 IEEE/CVF
                     Conf. on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018. 9185–9193. [doi: 10.1109/CVPR.2018.00957]
                 [16]   Carlini N, Wagner D. Towards evaluating the robustness of neural networks. In: Proc. of the 2017 IEEE Symp. on Security and Privacy.
                     San Jose: IEEE, 2017. 39–57. [doi: 10.1109/SP.2017.49]
                 [17]   Croce F, Hein M. Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks. In: Proc. of the 37th
                     Int’l Conf. on Machine Learning. New York: JMLR.org, 2020. 2206–2216. [doi: 10.5555/3524938.3525144]
                 [18]   Croce F, Hein M. Minimally distorted adversarial examples with a fast adaptive boundary attack. In: Proc. of the 37th Int’l Conf. on
                     Machine Learning. New York: JMLR.org, 2020. 2196–2205. [doi: 10.5555/3524938.3525143]
   44   45   46   47   48   49   50   51   52   53   54