Page 134 - 《软件学报》2026年第6期
P. 134

舒善富 等: 基于自适应知识蒸馏的代码大模型轻量化                                                       2453


                 [32]   Bai  YT,  Jones  A,  Ndousse  K,  et  al.  Training  a  helpful  and  harmless  assistant  with  reinforcement  learning  from  human  feedback.
                     arXiv:2204.05862, 2022.
                 [33]   Ouyang L, Wu J, Jiang X, Almeida D, Wainwright CL, Mishkin P, Zhang C, Agarwal S, Slama K, Ray A, Schulman J, Hilton J, Kelton
                     F, Miller L, Simens M, Askell A, Welinder P, Christiano P, Leike J, Lowe R. Training language models to follow instructions with
                     human feedback. In: Proc. of the 36th Int’l Conf. on Neural Information Processing Systems. New Orleans: Curran Associates Inc., 2022.
                     27730–27744.
                 [34]   Li R, Allal LB, Zi YT, et al. StarCoder: May the source be with you! arXiv.2305.06161, 2023.
                 [35]   Feng ZY, Guo DY, Tang DY, Duan N, Feng XC, Gong M, Shou LJ, Qin B, Liu T, Jiang DX, Zhou M. CodeBERT: A pre-trained model
                     for  programming  and  natural  languages.  In:  Proc.  of  the  2020  Conf.  on  Empirical  Methods  in  Natural  Language  Processing,  2020.
                     1536–1547. [doi: 10.18653/v1/2020.findings-emnlp.139]
                 [36]   Luo ZY, Xu C, Zhao P, Sun QF, Geng XB, Hu WX, Tao CY, Ma J, Lin QW, Jiang DX. WizardCoder: Empowering code large language
                     models with evol-instruct. In: Proc. of the 12th Int’l Conf. on Learning Representations (ICLR). Vienna, 2024. [doi: 10.48550/arXiv.2306.
                     08568]
                 [37]   Jain SM. Introduction to Transformers for NLP: With the Hugging Face Library and Models to Solve Problems. Berkeley: Apress, 2022.
                     [doi: 10.1007/978-1-4842-8844-3]
                 [38]   Touvron H, Martin L, Stone K, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288, 2023.
                 [39]   Mirzadeh SI, Farajtabar M, Li A, Levine N, Matsukawa A, Ghasemzadeh H. Improved knowledge distillation via teacher assistant. Proc.
                     of the AAAI Conf. on Artificial Intelligence, 2020, 34(4): 5191–5198. [doi: 10.1609/aaai.v34i04.5963]
                 [40]   Beyer L, Zhai XH, Royer A, Markeeva L, Anil R, Kolesnikov A. Knowledge distillation: A good teacher is patient and consistent. In:
                     Proc. of the 2022 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). New Orleans: IEEE, 2022. 10915–10924. [doi:
                     10.1109/CVPR52688.2022.01065]
                 [41]   Ko J, Kim S, Chen TY, Yun SY. DistiLLM: Towards streamlined distillation for large language models. In: Proc. of the 41st Int’l Conf.
                     on Machine Learning. 2024. 24872–24895. [doi: 10.48550/arXiv.2402.03898]
                 [42]   Kim  Y,  Rush  AM.  Sequence-level  knowledge  distillation.  In:  Proc.  of  the  2016  Conf.  on  Empirical  Methods  in  Natural  Language
                     Processing. Austin: ACL, 2016. 1317–1327.[doi: 10.48550/arXiv.1606.07947]
                 [43]   Tian YL, Krishnan D, Isola P. Contrastive representation distillation. In: Proc. of the 10th Int’l Conf. on Learning Representations. 2022.
                     [doi: 10.48550/arXiv.1910.10699]
                 [44]   Liu C, Bao XL, Zhang HY, Zhang N, Hu HB, Zhang XH, Yan M. Guiding ChatGPT for better code generation: An empirical study. In:
                     Proc. of the 2024 IEEE Int’l Conf. on Software Analysis, Evolution and Reengineering. Rovaniemi: IEEE, 2024. 102–113. [doi: 10.1109/
                     SANER60148.2024.00018]
                 [45]   Chia YK, Hong PF, Bing LD, Poria S. INSTRUCTEVAL: Towards holistic evaluation of instruction-tuned large language models. In:
                     Proc. of the 1st Edition of the Workshop on the Scaling Behavior of Large Language Models (SCALE-LLM 2024). St. Julian’s: ACL,
                     2024, 35–64. [doi: 10.48550/arXiv.2306.04757]
                 [46]   Minka T. Divergence measures and message passing. Research Report, Microsoft, 2005.
                 [47]   Nguyen  AT,  Tran  T,  Gal  Y,  Torr  PHS,  Baydin  AG.  KL  guided  domain  adaptation.  In:  Proc.  of  the  10th  Int’l  Conf.  on  Learning
                     Representations. 2022. [doi: 10.48550/arXiv.2106.07780]
                 [48]   Czarnecki WM, Pascanu R, Osindero S, Jayakumar SM, Świrszcz G, Jaderberg M. Distilling policy distillation. In: Proc. of the 22nd Int’l
                     Conf. on Artificial Intelligence and Statistics. Naha: PMLR, 2019. 1331–1340.
                 [49]   Chen M, Tworek J, Jun H, et al. Evaluating large language models trained on code. arXiv:2107.03374, 2021.
                 [50]   Austin J, Odena A, Nye M, Bosma M, Michalewski H, Dohan D, Jiang E, Cai C, Terry M, Le Q, Sutton C. Program synthesis with large
                     language models. arXiv:2108.07732, 2021.
                 [51]   Wen H, Zhu YH, Liu C, Ren XX, Du WW, Yan M. Fixing function-level code generation errors for foundation large language models.
                     arXiv:2409.00676, 2025.
                 [52]   Lee H, Park Y, Seo H, Kang M. Self-knowledge distillation via dropout. Computer Vision and Image Understanding, 2023, 233: 103720.
                     [doi: 10.1016/j.cviu.2023.103720]
                 [53]   Gunel B, Du JF, Conneau A, Stoyanov V. Supervised contrastive learning for pre-trained language model fine-tuning. In: Proc. of the 9th
                     Int’l Conf. on Learning Representations. 2021. [doi: 10.48550/arXiv.2011.01403]
                 [54]   Hu EJ, Shen YL, Wallis P, Allen-Zhu Z, Li YZ, Wang SA, Wang L, Chen WZ. LoRA: Low-rank adaptation of large language models. In:
                     Proc. of the 10th Int’l Conf. on Learning Representations. 2021. [doi: 10.48550/arXiv.2106.09685]
                 [55]   Wei YX, Wang Z, Liu JW, Ding YF, Zhang LM. Magicoder: Empowering code generation with OSS-instruct. In: Proc. of the 41st Int’l
   129   130   131   132   133   134   135   136   137   138   139