Page 203 - 《软件学报》2026年第2期
P. 203

682                                                        软件学报  2026  年第  37  卷第  2  期


                     for code infilling and synthesis. In: Proc. of the 11th Int’l Conf. on Learning Representations. Kigali: OpenReview.net, 2023.
                 [36]   Ahmad W, Chakraborty S, Ray B, Chang KW. Unified pre-training for program understanding and generation. In: Proc. of the 2021 Conf.
                     of  the  North  American  Chapter  of  the  Association  for  Computational  Linguistics:  Human  Language  Technologies.  ACL,  2021.
                     2655–2668. [doi: 10.18653/v1/2021.naacl-main.211]
                 [37]   Wang Y, Wang WS, Joty S, Hoi SCH. CodeT5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and
                     generation. In: Proc. of the 2021 Conf. on Empirical Methods in Natural Language Processing. Punta Cana: ACL, 2021. 8696–8708. [doi:
                     10.18653/v1/2021.emnlp-main.685]
                 [38]   Gao SZ, Gao CY, He YL, Zeng JC, Nie L, Xia X, Lyu M. Code structure guided Transformer for source code summarization. ACM
                     Trans. on Software Engineering and Methodology, 2023, 32(1): 23. [doi: 10.1145/3522674]
                 [39]   Nagaraj Y, Gupta U. AST-MHSA: Code summarization using multi-head self-attention. arXiv:2308.05646, 2023.
                 [40]   Shahbazi R, Fard F. APIContext2Com: Code comment generation by incorporating pre-defined API documentation. In: 2023 IEEE/ACM
                     31st Int’l Conf. on Program Comprehension. Melbourne: IEEE, 2023. 13–24. [doi: 10.1109/ICPC58990.2023.00012]
                 [41]   Li MC, Yu HQ, Fan GS, Zhou ZY, Huang ZJ. Enhancing code summarization with action word prediction. Neurocomputing, 2024, 563:
                     126777. [doi: 10.1016/j.neucom.2023.126777]
                 [42]   Sun WS, Fang CR, Chen YC, Zhang QJ, Tao GH, You YD, Han TX, Ge YF, Hu YL, Luo B, Chen ZY. An extractive-and-abstractive
                     framework  for  source  code  summarization.  ACM  Trans.  on  Software  Engineering  and  Methodology,  2024,  33(3):  75.  [doi:  10.1145/
                     3632742]
                 [43]   Wang Y, Le H, Gotmare A, Bui N, Li JN, Hoi S. CodeT5+: Open code large language models for code understanding and generation. In:
                     Proc. of the 2023 Conf. on Empirical Methods in Natural Language Processing. Singapore: ACL, 2023. 1069–1088. [doi: 10.18653/v1/
                     2023.emnlp-main.68]
                 [44]   Guo DY, Zhu QH, Yang DJ, Xie ZD, Dong K, Zhang WT, Chen GT, Bi X, Wu U, Li YK, Luo FL, Xiong YF, Liang WF. DeepSeek-
                     coder: When the large language model meets programming—The rise of code intelligence. arXiv:2401.14196, 2024.
                 [45]   Sun Y, Wang SH, Li YK, Feng SK, Chen XY, Zhang H, Tian X, Zhu DX, Tian H, Wu H. ERNIE: Enhanced representation through
                     knowledge integration. arXiv:1904.09223, 2019.
                 [46]   Clark K, Luong MT, Le QV, Manning CD. ELECTRA: Pre-training text encoders as discriminators rather than generators. arXiv:2003.
                     10555, 2020
                 [47]   Li  ZC,  Li  SS,  Zhou  GD.  Pre-trained  token-replaced  detection  model  as  few-shot  learner.  In:  Proc.  of  the  29th  Int’l  Conf.  on
                     Computational Linguistics. Gyeongju: ACL, 2022. 3274–3284.
                 [48]   Antoun W, Baly F, Hajj H. AraELECTRA: Pre-training text discriminators for Arabic language understanding. In: Proc. of the 6th Arabic
                     Natural Language Processing Workshop. Kyiv: ACL, 2021. 191–195.
                 [49]   He  PC,  Gao  JF,  Chen  WZ.  DeBERTaV3:  Improving  DeBERTa  using  ELECTRA-style  pre-training  with  gradient-disentangled
                     embedding sharing. In: Proc. of the 11th Int’l Conf. on Learning Representations. Kigali: OpenReivew.net, 2023.
                 [50]   Lachaux MA, Roziere B, Szafraniec M, Lample G. DOBF: A deobfuscation pre-training objective for programming languages. In: Proc.
                     of the 35th Int’l Conf. on Neural Information Processing Systems. Curran Associates Inc., 2021. 1147.
                 [51]   Jain  P,  Jain  A,  Zhang  TJ,  Abbeel  P,  Gonzalez  J,  Stoica  I.  Contrastive  code  representation  learning.  In:  Proc.  of  the  2021  Conf.  on
                     Empirical Methods in Natural Language Processing. Punta Cana: ACL, 2021. 5954–5971. [doi: 10.18653/v1/2021.emnlp-main.482]
                 [52]   Ding YRB, Buratti L, Pujar S, Morari A, Ray B, Chakraborty S. Contrastive learning for source code with structural and functional
                     properties. arXiv:2110.03868, 2021.
                 [53]   Sennrich R, Haddow B, Birch A. Neural machine translation of rare words with subword units. In: Proc. of the 54th Annual Meeting of
                     the Association for Computational Linguistics (Vol. 1: Long Papers). Berlin: ACL, 2016. 1715–1725. [doi: 10.18653/v1/P16-1162]
                 [54]   Radford A, Wu J, Child R, Luan D, Amodei D, Sutskever I. Language models are unsupervised multitask learners. OpenAI Blog, 2019,
                     1(8): 9.
                 [55]   Vinyals O, Fortunato M, Jaitly N. Pointer networks. In: Proc. of the 29th Int’l Conf. on Neural Information Processing Systems. Montreal:
                     MIT Press, 2015. 2692–2700.
                 [56]   Barone AVM, Sennrich R. A parallel corpus of Python functions and documentation strings for automated code documentation and code
                     generation. In: Proc. of the 8th Int’l Joint Conf. on Natural Language Processing (Vol. 2: Short Papers). Taipei: ACL, 2017. 314–319.
                 [57]   Husain H, Wu HH, Gazit T, Allamanis M, Brockschmidt M. CodeSearchNet challenge: Evaluating the state of semantic code search.
                     arXiv:1909.09436, 2019.
                 [58]   Papineni K, Roukos S, Ward T, Zhu WJ. BLEU: A method for automatic evaluation of machine translation. In: Proc. of the 40th Annual
                     Meeting of the Association for Computational Linguistics. Philadelphia: ACL, 2002. 311–318. [doi: 10.3115/1073083.1073135]
   198   199   200   201   202   203   204   205   206   207   208