Page 371 - 《软件学报》2026年第3期
P. 371
1334 软件学报 2026 年第 37 卷第 3 期
[7] Radford A, Wu J, Child R, Luan D, Amodei D, Sutskever I. Language models are unsupervised multitask learners. 2019. https://api.
semanticscholar.org/CorpusID:160025533
[8] Feng ZY, Guo DY, Tang DY, Duan N, Feng XC, Gong M, Shou LJ, Qin B, Liu T, Jiang DX, Zhou M. CodeBERT: A pre-trained model
for programming and natural languages. In: Findings of the Association for Computational Linguistics: EMNLP 2020. ACL, 2020.
1536–1547. [doi: 10.18653/v1/2020.findings-emnlp.139]
[9] Wang Y, Wang WS, Joty S, Hoi SCH. CodeT5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and
generation. In: Proc. of the 2021 Conf. on Empirical Methods in Natural Language Processing. Punta Cana: ACL, 2021. 8696–8708.
[doi: 10.18653/v1/2021.emnlp-main.685]
[10] Chen M, Tworek J, Jun H, et al. Evaluating large language models trained on code. arXiv:2107.03374, 2021.
[11] Zhang QJ, Fang CR, Zheng Y, Qian RX, Yu SC, Zhao Y, Zhou JY, Yang Y, Zheng T, Chen ZY. Improving retrieval-augmented deep
assertion generation via joint training. IEEE Trans. on Software Engineering, 2025, 51(4): 1232–1247. [doi: 10.1109/TSE.2025.
3545970]
[12] Yu XR, Li C, Pan MX, Li XD. DroidCoder: Enhanced android code completion with context-enriched retrieval-augmented generation.
In: Proc. of the 39th IEEE/ACM Int’l Conf. on Automated Software Engineering. Sacramento: ACM, 2024. 681–693. [doi: 10.1145/
3691620.3695063]
[13] Zhang QJ, Fang CR, Zheng Y, Zhang YX, Zhao Y, Huang RB, Zhou JY, Yang Y, Zheng T, Chen ZY. Improving deep assertion
generation via fine-tuning retrieval-augmented pre-trained language models. ACM Trans. on Software Engineering and Methodology,
2025, 34(7): 209. [doi: 10.1145/3721128]
[14] Su HJ, Jiang SY, Lai YH, Wu HY, Shi BA, Liu C, Liu Q, Yu T. EvoR: Evolving retrieval for code generation. In: Findings of the
Association for Computational Linguistics: EMNLP 2024. Miami: ACL, 2024. 2538–2554. [doi: 10.18653/v1/2024.findings-emnlp.143]
[15] Guo Q, Li XH, Xie XF, Liu SQ, Tang Z, Feng RT, Wang JJ, Ge JD, Bu L. FT2Ra: A fine-tuning-inspired approach to retrieval-
augmented code completion. In: Proc. of the 33rd ACM SIGSOFT Int’l Symp. on Software Testing and Analysis. Vienna: ACM, 2024.
313–324. [doi: 10.1145/3650212.3652130]
[16] Liang M, Xie XH, Zhang GH, Zheng XJ, Di P, Jiang W, Chen HW, Wang CP, Fan G. RepoGenix: Dual context-aided repository-level
code completion with language models. In: Proc. of the 39th IEEE/ACM Int’l Conf. on Automated Software Engineering. Sacramento:
ACM, 2024. 2466–2467. [doi: 10.1145/3691620.3695331]
[17] Wu SY, Xiong Y, Cui YF, Wu HL, Chen C, Yuan Y, Huang LM, Liu X, Kuo TW, Guan N, Xue CJ. Retrieval-augmented generation for
natural language processing: A survey. arXiv:2407.13193, 2024.
[18] Zhao PH, Zhang HL, Yu QH, Wang ZR, Geng YT, Fu FC, Yang L, Zhang WT, Jiang J, Cui B. Retrieval-augmented generation for AI-
generated content: A survey. arXiv:2402.19473, 2024.
[19] Gao YF, Xiong Y, Gao XY, Jia KX, Pan JL, Bi YX, Dai Y, Sun JW, Wang M, Wang HF. Retrieval-augmented generation for large
language models: A survey. arXiv:2312.10997, 2023.
[20] Fan WQ, Ding YJ, Ning LB, Wang SJ, Li HY, Yin DW, Chua TS, Li Q. A survey on RAG meeting LLMs: Towards retrieval-
augmented large language models. In: Proc. of the 30th ACM SIGKDD Conf. on Knowledge Discovery and Data Mining. Barcelona:
ACM, 2024. 6491–6501. [doi: 10.1145/3637528.3671470]
[21] Zhang QJ, Fang CR, Xie Y, Zhang YX, Yang Y, Sun WS, Yu SC, Chen ZY. A survey on large language models for software
engineering. arXiv:2312.15223, 2023.
[22] Petersen K, Vakkalanka S, Kuzniarz L. Guidelines for conducting systematic mapping studies in software engineering: An update.
Information and Software Technology, 2015, 64: 1–18. [doi: 10.1016/j.infsof.2015.03.007]
[23] Software Engineering Group. Guidelines for performing systematic literature reviews in software engineering. Technical Report, EBSE
2007-001. Durham: University of Durham, 2007. 1–65.
[24] Watson C, Cooper N, Palacio DN, Moran K, Poshyvanyk D. A systematic literature review on the use of deep learning in software
engineering research. ACM Trans. on Software Engineering and Methodology (TOSEM), 2022, 31(2): 32. [doi: 10.1145/3485275]
[25] Li J, Zhao YF, Li YM, Li G, Jin Z. ACECODER: An effective prompting technique specialized in code generation. ACM Trans. on
Software Engineering and Methodology, 2024, 33(8): 204. [doi: 10.1145/3675395]
[26] Li J, Li YM, Li G, Jin Z, Hao YY, Hu X. SkCoder: A sketch-based approach for automatic code generation. In: Proc. of the 45th
IEEE/ACM Int’l Conf. on Software Engineering (ICSE). Melbourne: IEEE, 2023. 2124–2135. [doi: 10.1109/ICSE48619.2023.00179]
[27] Zhang LH, Zhang HY, Wang C, Liang P. RAG-enhanced commit message generation. arXiv:2406.05514, 2024.
[28] Wang YL, Wang YL, Guo DY, Chen JC, Zhang RK, Ma YC, Zheng ZB. RLCoder: Reinforcement learning for repository-level code
completion. In: Proc. of the 47th IEEE/ACM Int’l Conf. on Software Engineering. Ottawa: IEEE, 2025. 1140–1152. [doi: 10.1109/

