Page 140 - 《软件学报》2026年第2期
P. 140
李重 等: 基于语义重排序的代码注释生成方法 619
IEEE/ACM Int’l Conf. on Automated Software Engineering. ACM, 2021. 746–757. [doi: 10.1145/3324884.3416546]
[13] Wei BL, Li YM, Li G, Xia X, Jin Z. Retrieve and refine: Exemplar-based neural comment generation. In: Proc. of the 35th IEEE/ACM
Int’l Conf. on Automated Software Engineering. ACM, 2021. 349–360. [doi: 10.1145/3324884.3416578]
[14] Zhang J, Wang X, Zhang HY, Sun HL, Liu XD. Retrieval-based neural source code summarization. In: Proc. of the 42nd ACM/IEEE Int’l
Conf. on Software Engineering. Seoul: ACM, 2020. 1385–1397. [doi: 10.1145/3377811.3380383]
[15] Li JA, Li YM, Li G, Hu X, Xia X, Jin Z. EditSum: A retrieve-and-edit framework for source code summarization. In: Proc. of the 36th
IEEE/ACM Int’l Conf. on Automated Software Engineering. Melbourne: IEEE, 2021. 155–166. [doi: 10.1109/ASE51524.2021.9678724]
[16] Zhu TW, Li Z, Pan MX, Shi CX, Zhang T, Pei Y, Li XD. Deep is better? An empirical comparison of information retrieval and deep
learning approaches to code summarization. ACM Trans. on Software Engineering and Methodology, 2024, 33(3): 67. [doi: 10.1145/
3631975]
[17] Hu X, Li G, Xia X, Lo D, Lu S, Jin Z. Summarizing source code with transferred API knowledge. In: Proc. of the 27th Int’l Joint Conf.
on Artificial Intelligence. Stockholm: IJCAI, 2018. 2269–2275. [doi: 10.24963/IJCAI.2018/314]
[18] Dify. 2024. https://docs.dify.ai/
[19] Chen T, Kornblith S, Norouzi M, Hinton G. A simple framework for contrastive learning of visual representations. In: Proc. of the 37th
Int’l Conf. on Machine Learning. JMLR.org, 2020. 1597–1607.
[20] Li Z, Pan MX, Pei Y, Zhang T, Wang LZ, Li XD. Empirically revisiting and enhancing automatic classification of bug and non-bug
issues. Frontiers of Computer Science, 2024, 18(5): 185207. [doi: 10.1007/S11704-023-2771-Z]
[21] Conneau A, Khandelwal K, Goyal N, Chaudhary V, Wenzek G, Guzmán F, Grave É, Ott M, Zettlemoyer L, Stoyanov V. Unsupervised
cross-lingual representation learning at scale. In: Proc. of the 58th Annual Meeting of the Association for Computational Linguistics.
ACL, 2020. 8440–8451. [doi: 10.18653/V1/2020.ACL-MAIN.747]
[22] Min BN, Ross H, Sulem E, Veyseh APB, Nguyen TH, Sainz O, Agirre E, Heintz I, Roth D. Recent advances in natural language
processing via large pre-trained language models: A survey. ACM Computing Surveys, 2024, 56(2): 30. [doi: 10.1145/3605943]
[23] Qiu XP, Sun TX, Xu YG, Shao YF, Dai N, Huang XJ. Pre-trained models for natural language processing: A survey. Science China
Technological Sciences, 2020, 63(10): 1872–1897. [doi: 10.1007/s11431-020-1647-3]
[24] Zheng ZB, Ning KW, Wang YL, Zhang JW, Zheng DW, Ye MX, Chen JC. A survey of large language models for code: Evolution,
benchmarking, and future trends. arXiv:2311.10372. 2024.
[25] Hugging Face. 2024. https://huggingface.co/
[26] Amershi S, Begel A, Bird C, DeLine R, Gall H, Kamar E, Nagappan N, Nushi B, Zimmermann T. Software engineering for machine
learning: A case study. In: Proc. of the 41st IEEE/ACM Int’l Conf. on Software Engineering: Software Engineering in Practice (ICSE-
SEIP). Montreal: IEEE, 2019. 291–300. [doi: 10.1109/ICSE-SEIP.2019.00042]
[27] Robertson SE, Walker S, Beaulieu M. Experimentation as a way of life: Okapi at TREC. Information Processing & Management, 2000,
36(1): 95–108. [doi: 10.1016/S0306-4573(99)00046-1]
[28] Liu SQ, Chen Y, Xie XF, Siow JK, Liu Y. Retrieval-augmented generation for code summarization via hybrid GNN. arXiv:2006.05405,
2021.
[29] Husain H, Wu HH, Gazit T, Allamanis M, Brockschmidt M. CodeSearchNet challenge: Evaluating the state of semantic code search.
arXiv:1909.09436, 2020.
[30] Han JF, Luo P, Wang XG. Deep self-learning from noisy labels. In: Proc. of the 2019 IEEE/CVF Int’l Conf. on Computer Vision. Seoul:
IEEE, 2019. 5137–5146. [doi: 10.1109/ICCV.2019.00524]
[31] Shi CX, Zhu TW, Zhang T, Pang J, Pan MX. Structural-semantics guided program simplification for understanding neural code
intelligence models. In: Proc. of the 14th Asia-Pacific Symp. on Internetware. Hangzhou: ACM, 2023. 1–11. [doi: 10.1145/3609437.
3609438]
[32] Papineni K, Roukos S, Ward T, Zhu WJ. BLEU: A method for automatic evaluation of machine translation. In: Proc. of the 40th Annual
Meeting on Association for Computational Linguistics. Philadelphia: ACL, 2002. 311–318. [doi: 10.3115/1073083.1073135]
[33] Wang Y, Wang WS, Joty S, Hoi SCH. CodeT5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and
generation. In: Proc. of the 2021 Conf. on Empirical Methods in Natural Language Processing. Punta Cana: ACL, 2021. 8696–8708.
[doi: 10.18653/V1/2021.EMNLP-MAIN.685]
[34] Barone AVM, Sennrich R. A parallel corpus of Python functions and documentation strings for automated code documentation and code
generation. In: Proc. of the 8th Int’l Joint Conf. on Natural Language Processing. Asian Federation of Natural Language Processing, 2017.
314–319.
[35] Lin C, Ouyang ZC, Zhuang JQ, Chen JQ, Li H, Wu RX. Improving code summarization with block-wise abstract syntax tree splitting. In:

