Page 134 - 《软件学报》2026年第6期
P. 134
舒善富 等: 基于自适应知识蒸馏的代码大模型轻量化 2453
[32] Bai YT, Jones A, Ndousse K, et al. Training a helpful and harmless assistant with reinforcement learning from human feedback.
arXiv:2204.05862, 2022.
[33] Ouyang L, Wu J, Jiang X, Almeida D, Wainwright CL, Mishkin P, Zhang C, Agarwal S, Slama K, Ray A, Schulman J, Hilton J, Kelton
F, Miller L, Simens M, Askell A, Welinder P, Christiano P, Leike J, Lowe R. Training language models to follow instructions with
human feedback. In: Proc. of the 36th Int’l Conf. on Neural Information Processing Systems. New Orleans: Curran Associates Inc., 2022.
27730–27744.
[34] Li R, Allal LB, Zi YT, et al. StarCoder: May the source be with you! arXiv.2305.06161, 2023.
[35] Feng ZY, Guo DY, Tang DY, Duan N, Feng XC, Gong M, Shou LJ, Qin B, Liu T, Jiang DX, Zhou M. CodeBERT: A pre-trained model
for programming and natural languages. In: Proc. of the 2020 Conf. on Empirical Methods in Natural Language Processing, 2020.
1536–1547. [doi: 10.18653/v1/2020.findings-emnlp.139]
[36] Luo ZY, Xu C, Zhao P, Sun QF, Geng XB, Hu WX, Tao CY, Ma J, Lin QW, Jiang DX. WizardCoder: Empowering code large language
models with evol-instruct. In: Proc. of the 12th Int’l Conf. on Learning Representations (ICLR). Vienna, 2024. [doi: 10.48550/arXiv.2306.
08568]
[37] Jain SM. Introduction to Transformers for NLP: With the Hugging Face Library and Models to Solve Problems. Berkeley: Apress, 2022.
[doi: 10.1007/978-1-4842-8844-3]
[38] Touvron H, Martin L, Stone K, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288, 2023.
[39] Mirzadeh SI, Farajtabar M, Li A, Levine N, Matsukawa A, Ghasemzadeh H. Improved knowledge distillation via teacher assistant. Proc.
of the AAAI Conf. on Artificial Intelligence, 2020, 34(4): 5191–5198. [doi: 10.1609/aaai.v34i04.5963]
[40] Beyer L, Zhai XH, Royer A, Markeeva L, Anil R, Kolesnikov A. Knowledge distillation: A good teacher is patient and consistent. In:
Proc. of the 2022 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). New Orleans: IEEE, 2022. 10915–10924. [doi:
10.1109/CVPR52688.2022.01065]
[41] Ko J, Kim S, Chen TY, Yun SY. DistiLLM: Towards streamlined distillation for large language models. In: Proc. of the 41st Int’l Conf.
on Machine Learning. 2024. 24872–24895. [doi: 10.48550/arXiv.2402.03898]
[42] Kim Y, Rush AM. Sequence-level knowledge distillation. In: Proc. of the 2016 Conf. on Empirical Methods in Natural Language
Processing. Austin: ACL, 2016. 1317–1327.[doi: 10.48550/arXiv.1606.07947]
[43] Tian YL, Krishnan D, Isola P. Contrastive representation distillation. In: Proc. of the 10th Int’l Conf. on Learning Representations. 2022.
[doi: 10.48550/arXiv.1910.10699]
[44] Liu C, Bao XL, Zhang HY, Zhang N, Hu HB, Zhang XH, Yan M. Guiding ChatGPT for better code generation: An empirical study. In:
Proc. of the 2024 IEEE Int’l Conf. on Software Analysis, Evolution and Reengineering. Rovaniemi: IEEE, 2024. 102–113. [doi: 10.1109/
SANER60148.2024.00018]
[45] Chia YK, Hong PF, Bing LD, Poria S. INSTRUCTEVAL: Towards holistic evaluation of instruction-tuned large language models. In:
Proc. of the 1st Edition of the Workshop on the Scaling Behavior of Large Language Models (SCALE-LLM 2024). St. Julian’s: ACL,
2024, 35–64. [doi: 10.48550/arXiv.2306.04757]
[46] Minka T. Divergence measures and message passing. Research Report, Microsoft, 2005.
[47] Nguyen AT, Tran T, Gal Y, Torr PHS, Baydin AG. KL guided domain adaptation. In: Proc. of the 10th Int’l Conf. on Learning
Representations. 2022. [doi: 10.48550/arXiv.2106.07780]
[48] Czarnecki WM, Pascanu R, Osindero S, Jayakumar SM, Świrszcz G, Jaderberg M. Distilling policy distillation. In: Proc. of the 22nd Int’l
Conf. on Artificial Intelligence and Statistics. Naha: PMLR, 2019. 1331–1340.
[49] Chen M, Tworek J, Jun H, et al. Evaluating large language models trained on code. arXiv:2107.03374, 2021.
[50] Austin J, Odena A, Nye M, Bosma M, Michalewski H, Dohan D, Jiang E, Cai C, Terry M, Le Q, Sutton C. Program synthesis with large
language models. arXiv:2108.07732, 2021.
[51] Wen H, Zhu YH, Liu C, Ren XX, Du WW, Yan M. Fixing function-level code generation errors for foundation large language models.
arXiv:2409.00676, 2025.
[52] Lee H, Park Y, Seo H, Kang M. Self-knowledge distillation via dropout. Computer Vision and Image Understanding, 2023, 233: 103720.
[doi: 10.1016/j.cviu.2023.103720]
[53] Gunel B, Du JF, Conneau A, Stoyanov V. Supervised contrastive learning for pre-trained language model fine-tuning. In: Proc. of the 9th
Int’l Conf. on Learning Representations. 2021. [doi: 10.48550/arXiv.2011.01403]
[54] Hu EJ, Shen YL, Wallis P, Allen-Zhu Z, Li YZ, Wang SA, Wang L, Chen WZ. LoRA: Low-rank adaptation of large language models. In:
Proc. of the 10th Int’l Conf. on Learning Representations. 2021. [doi: 10.48550/arXiv.2106.09685]
[55] Wei YX, Wang Z, Liu JW, Ding YF, Zhang LM. Magicoder: Empowering code generation with OSS-instruct. In: Proc. of the 41st Int’l

