Page 373 - 《软件学报》2026年第3期
P. 373
1336 软件学报 2026 年第 37 卷第 3 期
arXiv:2306.10998, 2023.
[48] Du XY, Zheng G, Wang KX, Zou Y, Wang YJ, Deng WT, Feng JY, Liu MW, Chen BH, Peng X, Ma T, Lou YL. Vul-RAG: Enhancing
LLM-based vulnerability detection via knowledge-level RAG. arXiv:2406.11147, 2024.
[49] Mackay A. Test suite augmentation using language models-applying RAG to improve robustness verification. In: Proc. of the 12th
European Congress on Embedded Real-time Systems (ERTS 2024). 2024. 1–11.
[50] Ouyang SR, Yu WH, Ma KX, Xiao ZL, Zhang ZH, Jia MZ, Han JW, Zhang HM, Yu D. RepoGraph: Enhancing AI software
engineering with repository-level code graph. arXiv:2410.14684, 2024.
[51] Zhang Z, Liu XY, Lin YZ, Gao X, Sun HL, Yuan Y. LLM-based unit test generation via property retrieval. arXiv:2410.13542, 2024.
[52] Bogin B, Gupta S, Clark P, Sabharwal A. Leveraging code to improve in-context learning for semantic parsing. In: Proc. of the 2024
Conf. of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Mexico City:
ACL, 2024. 4971–5012. [doi: 10.18653/v1/2024.naacl-long.279]
[53] Liu XY, Lan B, Hu ZY, Liu Y, Zhang ZC, Wang F, Shieh MQ, Zhou WM. CodexGraph: Bridging large language models and code
repositories via code graph databases. In: Proc. of the 2025 Conf. of the Nations of the Americas Chapter of the Association for
Computational Linguistics: Human Language Technologies. Albuquerque: ACL, 2025. 142–160. [doi: 10.18653/v1/2025.naacl-long.7]
[54] Guo DY, Zhu QH, Yang DJ, Xie ZD, Dong K, Zhang WT, Chen GT, Bi X, Wu Y, Li YK, Luo FL, Xiong YF, Liang WF. DeepSeek-
Coder: When the large language model meets programming—The rise of code intelligence. arXiv:2401.14196, 2024.
[55] Chen JK, Hu X, Li ZH, Gao CY, Xia X, Lo D. Code search is all you need? Improving code suggestions with code search. In: Proc. of
the 46th IEEE/ACM Int’l Conf. on Software Engineering. Lisbon: ACM, 2024. 73. [doi: 10.1145/3597503.3639085]
[56] Rozière B, Gehring J, Gloeckle F, et al. Code LLaMA: Open foundation models for code. arXiv:2308.12950, 2023.
[57] CodeGemma Team. CodeGemma: Open code models based on gemma. arXiv:2406.11409, 2024.
[58] Yang ZZ, Chen SR, Gao CY, Li ZH, Li G, Lyu MRT. Deep learning based code generation methods: Literature review. Ruan Jian Xue
Bao/Journal of Software, 2024, 35(2): 604–628 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/6981.htm [doi: 10.
13328/j.cnki.jos.006981]
[59] He PF, Wang SW, Chowdhury S, Chen TH. Evaluating the effectiveness and efficiency of demonstration retrievers in RAG for coding
tasks. arXiv:2410.09662, 2024.
[60] Daneshvar SS, Nong Y, Yang X, Wang SW, Cai HP. VulScribeR: Exploring RAG-based vulnerability augmentation with LLMs. ACM
Trans. on Software Engineering and Methodology, 2025. [doi: 10.1145/3760775]
[61] Kamiya T. A RAG method for source code inquiry tailored to long-context LLMs. arXiv:2404.06082, 2024.
[62] Mansourian D, Olsson A, Sönnerhed L. A generative AI approach to native iOS and Android code translation: With and without
retrieval-augmented generation (RAG). 2025. https://www.diva-portal.org/smash/record.jsf?pid=diva2%3A1879944&dswid=-5661
[63] Xu TT, Liu K, Xia X. Survey on automated vulnerability repair. Ruan Jian Xue Bao/Journal of Software, 2024, 35(1): 136–158 (in
Chinese with English abstract). http://www.jos.org.cn/1000-9825/6828.htm [doi: 10.13328/j.cnki.jos.006828]
[64] Dai HP, Sun CA, Jin H, Xiao MJ. State-of-the-art survey on fuzz testing for deep learning system. Ruan Jian Xue Bao/Journal of
Software, 2023, 34(11): 5008–5028 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/6679.htm [doi: 10.13328/j.cnki.
jos.006679]
[65] Bui TD, Luu-Van DT, Nguyen TP, Nguyen TT, Nguyen S, Vo HD. RAMBO: Enhancing RAG-based repository-level method body
completion. arXiv:2409.15204, 2024.
[66] Wu YX, He PF, Wang ZH, Wang SW, Tian Y, Chen TH. A comprehensive framework for evaluating API-oriented code generation in
large language models. arXiv:2409.15228, 2024.
[67] Koziolek H, Grüner S, Hark R, Ashiwal V, Linsbauer S, Eskandani N. LLM-based and retrieval-augmented control code generation. In:
Proc. of the 1st Int’l Workshop on Large Language Models for Code. Lisbon: ACM, 2024. 22–29. [doi: 10.1145/3643795.3648384]
[68] Ding YRB, Wang ZJ, Ahmad W, Ramanathan MK, Nallapati R, Bhatia P, Roth D, Xiang B. CoCoMIC: Code completion by jointly
modeling in-file and cross-file context. In: Proc. of the 2024 Joint Int’l Conf. on Computational Linguistics, Language Resources and
Evaluation. Torino: ACL, 2024. 3433–3445.
[69] Jin M, Shahriar S, Tufano M, Shi X, Lu S, Sundaresan N, Svyatkovskiy A. InferFix: End-to-end program repair with LLMs. In: Proc. of
the 31st ACM Joint European Software Engineering Conf. and Symp. on the Foundations of Software Engineering. San Francisco:
ACM, 2023. 1646–1656. [doi: 10.1145/3611643.3613892]
[70] Lu HZ, Liu ZX. Improving retrieval-augmented code comment generation by retrieving for generation. In: Proc. of the 2024 IEEE Int’l
Conf. on Software Maintenance and Evolution. Flagstaff: IEEE, 2024. 350–362. [doi: 10.1109/ICSME58944.2024.00040]
[71] Xu JJL, Cui Z, Zhao Y, Zhang X, He SL, He PJ. UniLog: Automatic logging via LLM and in-context learning. In: Proc. of the 46th

