Page 199 - 《软件学报》2026年第7期
P. 199
2884 软件学报 2026 年第 37 卷第 7 期
[15] Liu JW, Xia CS, Wang YY, Zhang LM. Is your code generated by ChatGPT really correct? Rigorous evaluation of large language models
for code generation. In: Proc. of the 37th Int’l Conf. on Neural Information Processing Systems. New Orleans: Curran Associates Inc.,
2023. 943.
[16] King JC. Symbolic execution and program testing. Communications of the ACM, 1976, 19(7): 385–394. [doi: 10.1145/360248.360252]
[17] Cadar C, Dunbar D, Engler D. KLEE: Unassisted and automatic generation of high-coverage tests for complex systems programs. In:
Proc. of the 8th USENIX Conf. on Operating Systems Design and Implementation. San Diego: USENIX Association, 2008. 209–224.
[18] Cadar C, Sen K. Symbolic execution for software testing: Three decades later. Communications of the ACM, 2013, 56(2): 82–90. [doi: 10.
1145/2408776.2408795]
[19] Cadar C, Nowack M. KLEE symbolic execution engine in 2019. Int’l Journal on Software Tools for Technology Transfer, 2021, 23(6):
867–870. [doi: 10.1007/s10009-020-00570-3]
[20] Iyer S, Konstas I, Cheung A, Zettlemoyer L. Mapping language to code in programmatic context. In: Proc. of the 2018 Conf. on
Empirical Methods in Natural Language Processing. Brussels: ACL, 2018. 1643–1652. [doi: 10.18653/v1/D18-1192]
[21] Li YJ, Choi D, Chung J, et al. Competition-level code generation with AlphaCode. Science, 2022, 378(6624): 1092–1097. [doi: 10.1126/
science.abq1158]
[22] Zhao WX, Zhou K, Li JY, Tang TY, Wang XL, Hou YP, Min YQ, Zhang BC, Zhang JJ, Dong ZC, Du YF, Yang C, Chen YS, Chen ZP,
Jiang JH, Ren RY, Li YF, Tang XY, Liu ZK, Liu PY, Nie JY, Wen JR. A survey of large language models. arXiv:2303.18223, 2023.
[23] OpenAI, Achiam J, Adler S, et al. GPT-4 technical report. arXiv:2303.08774, 2023.
[24] Anthropic. The Claude 3 model family: Opus, Sonnet, Haiku. 2024. https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc
618857627/Model_Card_Claude_3.pdf
[25] Rozière B, Gehring J, Gloeckle F, et al. Code Llama: Open foundation models for code. arXiv:2308.12950, 2023.
[26] Team G, Anil R, Borgeaud S, et al. Gemini: A family of highly capable multimodal models. arXiv:2312.11805, 2023.
[27] Team G, Georgiev P, Lei VI, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv:
2403.05530, 2024.
[28] DeepSeek-AI, Zhu QH, Guo DY, et al. DeepSeek-Coder-V2: Breaking the barrier of closed-source models in code intelligence.
arXiv:2406.11931, 2024.
[29] Du ZX, Qian YJ, Liu X, Ding M, Qiu JZ, Yang ZL, Tang J. GLM: General language model pretraining with autoregressive blank
infilling. In: Proc. of the 60th Annual Meeting of the Association for Computational Linguistics. Dublin: ACL, 2022. 320–335. [doi: 10.
18653/v1/2022.acl-long.26]
[30] Zeng AH, Liu X, Du ZX, Wang ZH, Lai HY, Ding M, Yang ZY, Xu YF, Zheng WD, Xia X, Tam W, Ma ZX, Xue YF, Zhai JD, Chen
WG, Liu ZY, Zhang P, Dong YX, Tang J. GLM-130B: An open bilingual pre-trained model. arXiv:2210.02414, 2022.
[31] Yu T, Zhang R, Yang K, Yasunaga M, Wang DX, Li ZF, Ma J, Li I, Yao QN, Roman S, Zhang ZL, Radev D. Spider: A large-scale
human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task. In: Proc. of the 2018 Conf. on Empirical
Methods in Natural Language Processing. Brussels: ACL, 2018. 3911–3921. [doi: 10.18653/v1/D18-1425]
[32] Cassano F, Gouwar J, Nguyen D, Nguyen S, Phipps-Costin L, Pinckney D, Yee MH, Zi YT, Anderson CJ, Feldman MQ, Guha A,
Greenberg M, Jangda A. MultiPL-E: A scalable and polyglot approach to benchmarking neural code generation. IEEE Trans. on Software
Engineering, 2023, 49(7): 3675–3691. [doi: 10.1109/TSE.2023.3267446]
[33] Miller BP, Fredriksen L, So B. An empirical study of the reliability of UNIX utilities. Communications of the ACM, 1990, 33(12): 32–44.
[doi: 10.1145/96267.96279]
[34] Holler C, Herzig K, Zeller A. Fuzzing with code fragments. In: Proc. of the 21st USENIX Conf. on Security Symp. Bellevue: USENIX
Association, 2012. 38.
[35] Cadar C, Godefroid P, Khurshid S, Păsăreanu CS, Sen K, Tillmann N, Visser W. Symbolic execution for software testing in practice:
Preliminary assessment. In: Proc. of the 33rd Int’l Conf. on Software Engineering. Honolulu: ACM, 2011. 1066–1071. [doi: 10.1145/
1985793.1985995]
[36] Bailey J, Nicholas C. Symbolic execution in practice: A survey of applications in vulnerability, malware, firmware, and protocol analysis.
arXiv:2508.06643, 2025.
[37] Vouvoutsis V, Casino F, Patsakis C. Beyond the sandbox: Leveraging symbolic execution for evasive malware classification. Computers
& Security, 2025, 149: 104193. [doi: 10.1016/j.cose.2024.104193]
[38] Chen JC, Shao ZZ, Yang S, Shen YM, Wang YL, Chen T, Shan ZY, Zheng ZB. NumScout: Unveiling numerical defects in smart
contracts using LLM-pruning symbolic execution. IEEE Trans. on Software Engineering, 2025, 51(5): 1538–1553. [doi: 10.1109/TSE.
2025.3555622]

