Page 232 - 《软件学报》2026年第3期
P. 232
田朝 等: CoDefense: 面向对抗性攻击的多粒度代码归一化防御方法 1195
semantic-preserving program transformations. Information and Software Technology, 2021, 135: 106552. [doi: 10.1016/j.infsof.2021.
106552]
[73] Pour MV, Li Z, Ma L, Hemmati H. A search-based testing framework for deep neural networks of source code embedding. In: Proc. of
the 14th IEEE Conf. on Software Testing, Verification and Validation. Porto de Galinhas: IEEE, 2021. 36–46. [doi: 10.1109/ICST49551.
2021.00016]
[74] Zhou Y, Zhang XQ, Shen JJ, Han TT, Chen TL, Gall H. Adversarial robustness of deep code comment generation. ACM Trans. on
Software Engineering and Methodology, 2022, 31(4): 60. [doi: 10.1145/3501256]
[75] Dong ZM, Hu Q, Guo YJ, Cordy M, Papadakis M, Zhang ZY. MixCode: Enhancing code classification by mixup-based data
augmentation. In: Proc. of the 2023 IEEE Int’l Conf. on Software Analysis, Evolution and Reengineering. Macao: IEEE, 2023. 379–390.
[doi: 10.1109/SANER56733.2023.00043]
[76] Tree-sitter. 2024. https://tree-sitter.github.io/tree-sitter/
[77] srcML. 2024. https://www.srcml.org/
[78] Wolf T, Debut L, Sanh V, Chaumond J, Delangue C, Moi A, Cistac P, Rault T, Louf R, Funtowicz M, Davison J, Shleifer S, von Platen
P, Ma C, Jernite Y, Plu J, Xu CW, Le Scao T, Gugger S, Drame M, Lhoest Q, Rush AM. Huggingface’s Transformers: State-of-the-art
natural language processing. arXiv:1910.03771, 2019.
[79] Rosner B, Glynn RJ, Lee MLT. The Wilcoxon signed rank test for paired comparisons of clustered data. Biometrics, 2006, 62(1):
185–192. [doi: 10.1111/j.1541-0420.2005.00389.x]
[80] Burkart N, Huber MF. A survey on the explainability of supervised machine learning. Journal of Artificial Intelligence Research, 2021,
70: 245–317. [doi: 10.1613/jair.1.12228]
[81] Tian Z, Shu HL, Wang D, Cao XJ, Kamei Y, Chen JJ. Large language models for equivalent mutant detection: How far are we? In:
Proc. of the 33rd ACM SIGSOFT Int’l Symp. on Software Testing and Analysis. Vienna: ACM, 2024. 1733–1745. [doi: 10.1145/
3650212.3680395]
[82] van der Maaten L, Hinton G. Visualizing data using t-SNE. Journal of Machine Learning Research, 2008, 9(11): 2579–2605.
[83] Li XY, Meng GZ, Liu SQ, Xiang L, Sun K, Chen K, Luo XP, Liu Y. Attribution-guided adversarial code prompt generation for code
completion models. In: Proc. of the 39th IEEE/ACM Int’l Conf. on Automated Software Engineering. Sacramento: ACM, 2024.
1460–1471. [doi: 10.1145/3691620.3695517]
[84] Zhang C, Wang ZF, Zhao RS, Mangal R, Fredrikson M, Jia LM, Păsăreanu CS. Attacks and defenses for large language models on
coding tasks. In: Proc. of the 39th IEEE/ACM Int’l Conf. on Automated Software Engineering. Sacramento: IEEE, 2024. 2268–2272.
[85] Yang Y, Yao H, Yang B. TAPI: Towards target-specific and adversarial prompt injection against code LLMs. arXiv:2407.09164, 2024.
[86] Guo DY, Zhu QH, Yang DJ, Xie ZD, Dong K, Zhang WT, Chen GT, Bi X, Wu Y, Li YK, Luo FL, Xiong YF, Liang WF. DeepSeek-
Coder: When the large language model meets programming—The rise of code intelligence. arXiv:2401.14196, 2024.
[87] Xiao Y, Lin Y, Beschastnikh I, Sun CS, Rosenblum D, Dong JS. Repairing failure-inducing inputs with input reflection. In: Proc. of the
37th IEEE/ACM Int’l Conf. on Automated Software Engineering. Rochester: ACM, 2022. 85. [doi: 10.1145/3551349.3556932]
[88] Karer HH, Soni PB. Dead code elimination technique in eclipse compiler for Java. In: Proc. of the 2015 Int’l Conf. on Control,
Instrumentation, Communication and Computational Technologies. Kumaracoil: IEEE, 2015. 275–278. [doi: 10.1109/ICCICCT.2015.
7475289]
[89] Jha A, Reddy CK. CodeAttack: Code-based adversarial attacks for pre-trained programming language models. In: Proc. of the 2023
AAAI Conf. on Artificial Intelligence. Washington: AAAI Press, 2023. 14892–14900. [doi: 10.1609/aaai.v37i12.26739]
[90] Liu DX, Zhang SK. ALANCA: Active learning guided adversarial attacks for code comprehension on diverse pre-trained and large
language models. In: Proc. of the 2024 IEEE Int’l Conf. on Software Analysis, Evolution and Reengineering. Rovaniemi: IEEE, 2024.
602–613. [doi: 10.1109/SANER60148.2024.00067]
[91] Bui NDQ, Yu YJ, Jiang LX. TreeCaps: Tree-based capsule networks for source code processing. In: Proc. of the 2021 AAAI Conf. on
Artificial Intelligence. AAAI Press, 2021. 30–38. [doi: 10.1609/aaai.v35i1.16074]
[92] Wang X, Wang YS, Mi F, Zhou PY, Wan Y, Liu X, Li L, Wu H, Liu J, Jiang X. SynCoBERT: Syntax-guided multi-modal contrastive
pre-training for code representation. arXiv:2108.04556, 2021.
[93] Liu SQ, Wu BZ, Xie XF, Meng GZ, Liu Y. ContraBERT: Enhancing code pre-trained models via contrastive learning. In: Proc. of the
45th IEEE/ACM Int’l Conf. on Software Engineering. Melbourne: IEEE, 2023. 2476–2487. [doi: 10.1109/ICSE48619.2023.00207]
[94] Wang X, Wu Q, Zhang HY, Lyu C, Jiang X, Zheng ZR. HELoC: Hierarchical contrastive learning of source code representation. In:
Proc. of the 30th IEEE/ACM Int’l Conf. on Program Comprehension. Pittsburgh: IEEE, 2022. 354–365. [doi: 10.1145/3524610.
3527896]

