Page 230 - 《软件学报》2026年第3期
P. 230

田朝 等: CoDefense: 面向对抗性攻击的多粒度代码归一化防御方法                                           1193


                      Association for Computational Linguistics. ACL, 2020. 7871–7880. [doi: 10.18653/v1/2020.acl-main.703]
                 [27]   Raffel C, Shazeer N, Roberts A, Lee K, Narang S, Matena M, Zhou YQ, Li W, Liu PJ. Exploring the limits of transfer learning with a
                      unified text-to-text Transformer. The Journal of Machine Learning Research, 2020, 21(1): 140.
                 [28]   Wang SQ, Li Z, Qian HF, Yang CH, Wang ZJ, Shang MY, Kumar V, Tan S, Ray B, Bhatia P, Nallapati R, Ramanathan MK, Roth D,
                      Xiang  B.  Recode:  Robustness  evaluation  of  code  generation  models.  In:  Proc.  of  the  61st  Annual  Meeting  of  the  Association  for
                      Computational Linguistics. Toronto: ACL, 2022. 13818–13843. [doi: 10.18653/v1/2023.acl-long.773]
                 [29]   Karampatsis RM, Babii H, Robbes R, Sutton C, Janes A. Big code != big vocabulary: Open-vocabulary models for source code. In:
                      Proc. of the 42nd ACM/IEEE Int’l Conf. on Software Engineering. Seoul: ACM, 2020. 1073–1085. [doi: 10.1145/3377811.3380342]
                 [30]   Yefet  N,  Alon  U,  Yahav  E.  Adversarial  examples  for  models  of  code.  Proc.  of  the  ACM  on  Programming  Languages,  2020,
                      4(OOPSLA): 162. [doi: 10.1145/3428230]
                 [31]   Tian  Z,  Chen  JJ,  Zhang  XY.  On-the-fly  improving  performance  of  deep  code  models  via  input  denoising.  In:  Proc.  of  the  38th
                      IEEE/ACM Int’l Conf. on Automated Software Engineering. Luxembourg: IEEE, 2023. 560–572. [doi: 10.1109/ASE56229.2023.00166]
                 [32]   Brown WH, Malveau RC, McCormick HWS, Mowbray TJ. AntiPatterns: Refactoring Software, Architectures, and Projects in Crisis.
                      New York: John Wiley & Sons, Inc., 1998.
                 [33]   Mantyla M, Vanhanen J, Lassenius C. A taxonomy and an initial empirical study of bad smells in code. In: Proc. of the 2003 Int’l Conf.
                      on Software Maintenance. Amsterdam: IEEE, 2003. 381–384. [doi: 10.1109/ICSM.2003.1235447]
                 [34]   Wake WC. Refactoring Workbook. Boston: Addison-Wesley Longman, 2003.
                 [35]   Martin RC. Clean code: A Handbook of Agile Software Craftsmanship. Upper Saddle River: Pearson, 2008.
                 [36]   Romano  S,  Vendome  C,  Scanniello  G,  Poshyvanyk  D.  A  multi-study  investigation  into  dead  code.  IEEE  Trans.  on  Software
                      Engineering, 2020, 46(1): 71–99. [doi: 10.1109/TSE.2018.2842781]
                 [37]   Nguyen  TD,  Zhou  Y,  Le  XBD,  Thongtanunam  P,  Lo  D.  Adversarial  attacks  on  code  models  with  discriminative  graph  patterns.
                      arXiv:2308.11161, 2023.
                 [38]   Bui  NDQ,  Yu  YJ,  Jiang  LX.  Self-supervised  contrastive  learning  for  code  retrieval  and  summarization  via  semantic-preserving
                      transformations. In: Proc. of the 44th Int’l ACM SIGIR Conf. on Research and Development in Information Retrieval. ACM, 2021.
                      511–521. [doi: 10.1145/3404835.3462840]
                 [39]   Srikant  S,  Liu  SJ,  Mitrovska  T,  Chang  SY,  Fan  QF,  Zhang  GY,  O’Reilly  UM.  Generating  adversarial  computer  programs  using
                      optimized obfuscations. arXiv:2103.11882, 2021.
                 [40]   Jia JH, Srikant S, Mitrovska T, Gan C, Chang SY, Liu SJ. ClawSAT: Towards both robust and accurate code models. In: Proc. of the
                      2023 IEEE Int’l Conf. on Software Analysis, Evolution and Reengineering. Macao: IEEE, 2023. 212–223. [doi: 10.1109/SANER56733.
                      2023.00029]
                 [41]   Li J, Li Z, Zhang HZ, Li G, Jin Z, Hu X, Xia X. Poison attack and defense on deep source code processing models. arXiv:2210.17029, 2022.
                 [42]   Bielik P, Vechev M. Adversarial robustness for code. In: Proc. of the 37th Int’l Conf. on Machine Learning. JMLR.org, 2020. 896–907.
                      [doi: 10.48550/arXiv.2002.04694]
                 [43]   Zhuo TY, Yang Z, Sun ZS, Wang YF, Li L, Du XN, Xing ZC, Lo D. Data augmentation approaches for source code models: A survey.
                      arXiv:2305.19915, 2023.
                 [44]   Mercuri V, Saletta M, Ferretti C. Evolutionary approaches for adversarial attacks on neural source code classifiers. Algorithms, 2023,
                      16(10): 478. [doi: 10.3390/a16100478]
                 [45]   Henkel J, Ramakrishnan G, Wang Z, Albarghouthi A, Jha S, Reps T. Semantic robustness of models of source code. In: Proc. of the
                      2022  IEEE  Int’l  Conf.  on  Software  Analysis,  Evolution  and  Reengineering.  Honolulu:  IEEE,  2022.  526–537.  [doi:  10.1109/
                      SANER53432.2022.00070]
                 [46]   Gopstein D, Iannacone J, Yan Y, DeLong L, Zhuang YY, Yeh MKC, Cappos J. Understanding misunderstandings in source code. In:
                      Proc. of the 11th Joint Meeting on Foundations of Software Engineering. Paderborn: ACM, 2017. 129–139. [doi: 10.1145/3106237.
                      3106264]
                 [47]   Li  Z,  Chen  GQ,  Chen  C,  Zou  YY,  Xu  SH.  RoPGen:  Towards  robust  code  authorship  attribution  via  automatic  coding  style
                      transformation. In: Proc. of the 44th IEEE/ACM Int’l Conf. on Software Engineering. Pittsburgh: IEEE, 2022. 1906–1918. [doi: 10.1145/
                      3510003.3510181]
                 [48]   Chakraborty S, Ahmed T, Ding Y, Devanbu PT, Ray B. Natgen: Generative pre-training by “naturalizing” source code. In: Proc. of the
                      30th ACM Joint European Software Engineering Conf. and Symp. on the Foundations of Software Engineering. Singapore: ACM, 2022.
                      18–30. [doi: 10.1145/3540250.3549162]
                 [49]   Li Z, Zhang RQ, Zou DQ, Wang N, Li YT, Xu SH. Robin: A novel method to produce robust interpreters for deep learning-based code
   225   226   227   228   229   230   231   232   233   234   235