Page 272 - 《软件学报》2026年第4期
P. 272
汪莹 等: 基于大语言模型的故障复现测试用例生成方法 1713
[14] Baars A, Harman M, Hassoun Y, Lakhotia K, McMinn P, Tonella P, Vos T. Symbolic search-based testing. In: Proc. of the 26th
IEEE/ACM Int’l Conf. on Automated Software Engineering. Lawrence: IEEE, 2011. 53–62. [doi: 10.1109/ASE.2011.6100119]
[15] Blasi A, Gorla A, Ernst MD, Pezzè M. Call me maybe: Using NLP to automatically generate unit test cases respecting temporal
constraints. In: Proc. of the 37th IEEE/ACM Int’l Conf. on Automated Software Engineering. Rochester: IEEE, 2022. 19. [doi: 10.1145/
3551349.3556961]
[16] DeMilli RA, Offutt AJ. Constraint-based automatic test data generation. IEEE Trans. on Software Engineering, 1991, 17(9): 900–910.
[doi: 10.1109/32.92910]
[17] Pacheco C, Lahiri SK, Ernst MD, Ball T. Feedback-directed random test generation. In: Proc. of the 29th Int’l Conf. on Software
Engineering. Minneapolis: IEEE, 2007. 75–84. [doi: 10.1109/ICSE.2007.37]
[18] Dinella E, Ryan G, Mytkowicz T, Lahiri SK. TOGA: A neural method for test oracle generation. In: Proc. of the 44th Int’l Conf. on
Software Engineering. Pittsburgh: ACM, 2022. 2130–2141. [doi: 10.1145/3510003.3510141]
[19] Schäfer M, Nadi S, Eghbali A, Tip F. An empirical evaluation of using large language models for automated unit test generation. IEEE
Trans. on Software Engineering, 2024, 50(1): 85–105. [doi: 10.1109/TSE.2023.3334955]
[20] Yuan ZQ, Liu MW, Ding SJ, Wang KX, Chen YX, Peng X, Lou YL. Evaluating and improving ChatGPT for unit test generation. Proc.
of the ACM on Software Engineering, 2024, 1(FSE): 76. [doi: 10.1145/3660783]
[21] OpenDevin: Code less, make more. 2024. https://github.com/OpenDevin/OpenDevin/
[22] Chen D, Lin SX, Zeng MH, Zan DG, Wang JG, Cheshkov A, Sun J, Yu H, Dong GL, Aliev A, Wang J, Cheng X, Liang GT, Ma YC,
Bian P, Xie T, Wang QX. CodeR: Issue resolving with multi-agent and task graphs. arXiv:2406.01304, 2024.
[23] OpenAI. ChatGPT. 2024. https://openai.com/index/chatgpt/
[24] Li R, Allal LB, Zi YT, et al. StarCoder: May the source be with you! arXiv:2305.06161, 2023.
[25] Luo ZY, Xu C, Zhao P, Sun QF, Geng XB, Hu WX, Tao CY, Ma J, Lin QW, Jiang DX. WizardCoder: Empowering code large language
models with evol-instruct. In: Proc. of the 12th Int’l Conf. on Learning Representations. Vienna: OpenReview.net, 2024.
[26] Fried D, Aghajanyan A, Lin J, Wang SD, Wallace E, Shi F, Zhong RQ, Yih S, Zettlemoyer L, Lewis M. InCoder: A generative model for
code infilling and synthesis. In: Proc. of the 11th Int’l Conf. on Learning Representations. Kigali: OpenReview.net, 2023.
[27] Jin M, Shahriar S, Tufano M, Shi X, Lu S, Sundaresan N, Svyatkovskiy A. InferFix: End-to-end program repair with LLMs. In: Proc. of
the 31st ACM Joint European Software Engineering Conf. and Symp. on the Foundations of Software Engineering. San Francisco: ACM,
2023. 1646–1656. [doi: 10.1145/3611643.3613892]
[28] Xia CS, Zhang LM. Automated program repair via conversation: Fixing 162 out of 337 bugs for $0.42 each using ChatGPT. In: Proc. of
the 33rd ACM SIGSOFT Int’l Symp. on Software Testing and Analysis. Vienna: ACM, 2024. 819–831. [doi: 10.1145/3650212.3680323]
[29] Lewis P, Perez E, Piktus A, Petroni F, Karpukhin V, Goyal N, Küttler H, Lewis M, Yih WT, Rocktäschel T, Riedel S, Kiela D. Retrieval-
augmented generation for knowledge-intensive NLP tasks. In: Proc. of the 34th Int’l Conf. on Neural Information Processing Systems.
Vancouver: Curran Associates Inc., 2020. 793.
[30] Zhou SY, Alon U, Xu FF, Jiang ZB, Neubig G. DocPrompting: Generating code by retrieving the docs. In: Proc. of the 11th Int’l Conf.
on Learning Representations. Kigali: OpenReview.net, 2023.
[31] Lu S, Duan N, Han H, Guo DY, Hwang SW, Svyatkovskiy A. ReACC: A retrieval-augmented code completion framework. In: Proc. of
the 60th Annual Meeting of the Association for Computational Linguistics. Dublin: ACL, 2022. 6227–6240. [doi: 10.18653/v1/2022.acl-
long.431]
[32] Zhang FJ, Chen B, Zhang Y, Keung J, Liu J, Zan DG, Mao Y, Lou JG, Chen WZ. RepoCoder: Repository-level code completion through
iterative retrieval and generation. In: Proc. of the 2023 Conf. on Empirical Methods in Natural Language Processing. Singapore: ACL,
2023. 2471–2484. [doi: 10.18653/v1/2023.emnlp-main.151]
[33] Luan SF, Yang D, Barnaby C, Sen K, Chandra S. Aroma: Code recommendation via structural code search. Proc. of the ACM on
Programming Languages, 2019, 3(OOPSLA): 152. [doi: 10.1145/3360578]
[34] Wei J, Wang XZ, Schuurmans D, Bosma M, Ichter B, Xia F, Chi EH, Le QV, Zhou D. Chain-of-thought prompting elicits reasoning in
large language models. In: Proc. of the 36th Int’l Conf. on Neural Information Processing Systems. New Orleans: Curran Associates Inc.,
2022. 1800.
[35] White J, Hays S, Fu QC, Spencer-Smith J, Schmidt DC. ChatGPT prompt patterns for improving code quality, refactoring, requirements
elicitation, and software design. In: Nguyen-Duc A, Abrahamsson P, Khomh F, eds. Generative AI for Effective Software Development.
Cham: Springer, 2024. 71–108. [doi: 10.1007/978-3-031-55642-5_4]
[36] Xu BF, Yang A, Lin JY, Wang Q, Zhou C, Zhang YD, Mao ZD. ExpertPrompting: Instructing large language models to be distinguished
experts. arXiv:2305.14688, 2025.

