Page 198 - 《软件学报》2026年第7期
P. 198
赵祖威 等: 软件供应链安全中 LLM 生成代码逻辑性缺陷检测 2883
开发中, 仅会增加开发人员在缺陷验证与修复阶段对缺陷真实性进行人工确认的时间. 换言之, 本文方法仅会引入
约 7.2% 的误报, 而该误报所带来的额外确认时间仍处于可接受范围内, 不会对软件供应链的可靠性与安全性造成
实质影响.
5 总 结
本文提出一种融合符号执行的 LLM 生成代码缺陷检测方法, 用于检测基础软件供应链中 LLM 生成代码的
功能正确性, 保障基于 LLM 构造的智能化基础软件的安全性. 该方法利用符号执行从程序中提取精确条件作为符
号约束, 并结合 SMT 求解器对这些约束进行求解, 有效弥补了传统基于模糊测试的方法在深度逻辑推理方面的不
足. 因此, 本文方法能够生成补充性的边界测试用例, 有效覆盖供应链中生成代码的关键路径, 进而提升对 LLM 生
成代码的缺陷检测能力. 我们在 LMSYS Chatbot Arena 中排名前 11 的主流 LLM 上进行了实验评估, 实验结果表
明, 与现有方法相比, 本文方法在测试通过率方面降低了 3.99%–18.98%, 在代码覆盖率方面提升了 3.31%–8.19%,
这些结果体现了本文方法在软件供应链缺陷检测方面的有效性.
References
[1] Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I. Attention is all you need. In: Proc. of the
31st Int’l Conf. on Neural Information Processing Systems. Long Beach: Curran Associates Inc., 2017. 6000–6010.
[2] Liu YH, Ott M, Goyal N, Du JF, Joshi M, Chen DQ, Levy O, Lewis M, Zettlemoyer L, Stoyanov V. RoBERTa: A robustly optimized
BERT pretraining approach. arXiv:1907.11692, 2019.
[3] Brown TB, Mann B, Ryder N, et al. Language models are few-shot learners. In: Proc. of the 34th Int’l Conf. on Neural Information
Processing Systems. Vancouver: Curran Associates Inc., 2020. 159.
[4] Raffel C, Shazeer N, Roberts A, Lee K, Narang S, Matena M, Zhou YQ, Li W, Liu PJ. Exploring the limits of transfer learning with a
unified text-to-text Transformer. Journal of Machine Learning Research, 2020, 21(1): 140.
[5] Gulwani S, Polozov O, Singh R. Program synthesis. Foundations and Trends in Programming Languages, 2017, 4(1–2): 1–119. [doi: 10.
1561/2500000010]
[6] Austin J, Odena A, Nye M, Bosma M, Michalewski H, Dohan D, Jiang E, Cai C, Terry M, Le Q, Sutton C. Program synthesis with large
language models. arXiv:2108.07732, 2021.
[7] Chen M, Tworek J, Jun H, et al. Evaluating large language models trained on code. arXiv:2107.03374, 2021.
[8] Nijkamp E, Hayashi H, Xiong CM, Savarese S, Zhou YB. CodeGen2: Lessons for training LLMs on programming and natural languages.
arXiv:2305.02309, 2023.
[9] Rasnayaka S, Wang GL, Shariffdeen R, Iyer GN. An empirical study on usage and perceptions of LLMs in a software engineering
project. In: Proc. of the 1st Int’l Workshop on Large Language Models for Code. Lisbon: ACM, 2024. 111–118. [doi: 10.1145/3643795.
3648379]
[10] Williams L, Benedetti G, Hamer S, Paramitha R, Rahman I, Tamanna M, Tystahl G, Zahan N, Morrison P, Acar Y, Cukier M, Kästner C,
Kapravelos A, Wermke D, Enck W. Research directions in software supply chain security. ACM Trans. on Software Engineering and
Methodology, 2025, 34(5): 146. [doi: 10.1145/3714464]
[11] Feng ZY, Guo DY, Tang DY, Duan N, Feng XC, Gong M, Shou LJ, Qin B, Liu T, Jiang DX, Zhou M. CodeBERT: A pre-trained model
for programming and natural languages. In: Findings of the Association for Computational Linguistics: EMNLP 2020. ACL, 2020.
1536–1547. [doi: 10.18653/v1/2020.findings-emnlp.139]
[12] Imtiaz N, Thorn S, Williams L. A comparative study of vulnerability reporting by software composition analysis tools. In: Proc. of the
15th ACM/IEEE Int’l Symp. on Empirical Software Engineering and Measurement. Bari: ACM, 2021. 5. [doi: 10.1145/3475716.
3475769]
[13] Liu JQ, Tian X, Shu YQ, Zhu XX, Liu YL, Liu QX. A review of security research in the development stage of software supply chain
enhanced by large models. Journal of Cyber Security, 2024, 9(5): 87–109 (in Chinese with English abstract). [doi: 10.19363/J.cnki.cn10-
1380/tn.2024.09.11]
[14] Zheng QK, Xia X, Zou X, Dong YX, Wang S, Xue YF, Shen L, Wang ZH, Wang AD, Li Y, Su T, Yang ZL, Tang J. CodeGeeX: A pre-
trained model for code generation with multilingual benchmarking on HumanEval-X. In: Proc. of the 29th ACM SIGKDD Conf. on
Knowledge Discovery and Data Mining. Long Beach: ACM, 2023. 5673–5684. [doi: 10.1145/3580305.3599790]

