Page 372 - 《软件学报》2026年第3期
P. 372
张犬俊 等: 检索增强生成在软件工程中的应用综述 1335
ICSE55347.2025.00014]
[29] Parvez R, Ahmad W, Chakraborty S, Ray B, Chang KW. Retrieval augmented code generation and summarization. In: Findings of the
Association for Computational Linguistics: EMNLP 2021. Punta Cana: ACL, 2021. 2719–2734. [doi: 10.18653/v1/2021.findings-emnlp.
232]
[30] Gou QW, Dong YW, Wu YJ, Ke Q. RRGcode: Deep hierarchical search-based code generation. Journal of Systems and Software, 2024,
211: 111982. [doi: 10.1016/j.jss.2024.111982]
[31] Nan LY, Zhao YL, Zou WJ, Ri N, Tae J, Zhang E, Cohan A, Radev D. Enhancing text-to-SQL capabilities of large language models: A
study on prompt design strategies. In: Findings of the Association for Computational Linguistics: EMNLP 2023. Singapore: ACL, 2023.
14935–14956. [doi: 10.18653/v1/2023.findings-emnlp.996]
[32] Zhao JJ, Chen X, Yang G, Shen YH. Automatic smart contract comment generation via large language models and in-context learning.
Information and Software Technology, 2024, 168: 107405. [doi: 10.1016/j.infsof.2024.107405]
[33] Yu C, Yang G, Chen X, Liu K, Zhou YL. BashExplainer: Retrieval-augmented bash code comment generation based on fine-tuned
CodeBERT. In: Proc. of the 2022 IEEE Int’l Conf. on Software Maintenance and Evolution. Limassol: IEEE, 2022. 82–93. [doi: 10.
1109/ICSME55016.2022.00016]
[34] Guo DY, Ren S, Lu S, Feng ZY, Tang DY, Liu SJ, Zhou L, Duan N, Svyatkovskiy A, Fu SY, Tufano M, Deng SK, Clement CB, Drain
D, Sundaresan N, Yin J, Jiang DX, Zhou M. GraphCodeBERT: Pre-training code representations with data flow. In: Proc. of the 9th Int’l
Conf. on Learning Representations. OpenReview.net, 2021.
[35] Zhang FJ, Chen B, Zhang Y, Keung J, Liu J, Zan DG, Mao Y, Lou JG, Chen WZ. RepoCoder: Repository-level code completion
through iterative retrieval and generation. In: Proc. of the 2023 Conf. on Empirical Methods in Natural Language Processing. Singapore:
ACL, 2023. 2471–2484. [doi: 10.18653/v1/2023.emnlp-main.151]
[36] Li LX, Liang B, Chen L, Zhang XF. Cross-modal retrieval-enhanced code summarization based on joint learning for retrieval and
generation. Information and Software Technology, 2024, 175: 107527. [doi: 10.1016/j.infsof.2024.107527]
[37] Tang Z, Ge JD, Liu SQ, Zhu TW, Xu TT, Huang LG, Luo B. Domain adaptive code completion via language models and decoupled
domain databases. In: Proc. of the 38th IEEE/ACM Int’l Conf. on Automated Software Engineering. Luxembourg: IEEE, 2023.
421–433. [doi: 10.1109/ASE56229.2023.00076]
[38] Eghbali A, Pradel M. De-Hallucinator: Iterative grounding for LLM-based code completion. arXiv:2401.01701v1, 2024.
[39] Gong LN, Zhou YR, Qiao Y, Jiang SJ, Wei MQ, Huang ZQ. Research progress of pre-trained model in software engineering. Ruan Jian
Xue Bao/Journal of Software, 2025, 36(1): 1–26 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/7143.htm [doi: 10.
13328/j.cnki.jos.007143]
[40] Gao XY, Xiong Y, Wang DZ, Guan ZH, Shi ZJ, Wang HF, Li SS. Preference-guided refactored tuning for retrieval augmented code
generation. In: Proc. of the 39th IEEE/ACM Int’l Conf. on Automated Software Engineering. Sacramento: ACM, 2024. 65–77. [doi: 10.
1145/3691620.3694987]
[41] Wang WS, Wang Y, Joty S, Hoi SCH. RAP-Gen: Retrieval-augmented patch generation with CodeT5 for automatic program repair. In:
Proc. of the 31st ACM Joint European Software Engineering Conf. and Symp. on the Foundations of Software Engineering. San
Francisco: ACM, 2023. 146–158. [doi: 10.1145/3611643.3616256]
[42] Lu S, Duan N, Han H, Guo DY, Hwang SW, Svyatkovskiy A. ReACC: A retrieval-augmented code completion framework. In: Proc. of
the 60th Annual Meeting of the Association for Computational Linguistics. Dublin: ACL, 2022. 6227–6240. [doi: 10.18653/v1/2022.acl-
long.431]
[43] Wang Y, Le H, Gotmare A, Bui N, Li JN, Hoi S. CodeT5+: Open code large language models for code understanding and generation.
In: Proc. of the 2023 Conf. on Empirical Methods in Natural Language Processing. Singapore: ACL, 2023. 1069–1088. [doi: 10.18653/
v1/2023.emnlp-main.68]
[44] Liang M, Xie XH, Zhang GH, Zheng XJ, Di P, Jiang W, Chen HW, Wang CP, Fan G. REPOFUSE: Repository-level code completion
with fused dual context. arXiv:2402.14323, 2024.
[45] Liu MW, Yang TY, Lou YL, Du XY, Wang Y, Peng X. CodeGen4Libs: A two-stage approach for library-oriented code generation. In:
Proc. of the 38th IEEE/ACM Int’l Conf. on Automated Software Engineering. Luxembourg: IEEE, 2023. 434–445. [doi: 10.1109/
ASE56229.2023.00159]
[46] Wang ZZ, Asai A, Yu XV, Xu FF, Xie YQ, Neubig G, Fried D. CodeRAG-Bench: Can retrieval augment code generation? In: Findings
of the Association for Computational Linguistics: NAACL 2025. Albuquerque: ACL, 2024. 3199–3214. [doi: 10.18653/v1/2025.
findings-naacl.176]
[47] Shrivastava D, Kocetkov D, de Vries H, Bahdanau D, Scholak T. RepoFusion: Training code models to understand your repository.

