Page 133 - 《软件学报》2026年第6期
P. 133
2452 软件学报 2026 年第 37 卷第 6 期
evaluation and domain application. AI Open, 2024, 5: 155–180. [doi: 10.1016/j.aiopen.2024.08.001]
[9] Gu YX, Dong L, Wei FR, Huang ML. MiniLLM: Knowledge distillation of large language models. In: Proc. of the 12th Int’l Conf. on
Learning Representations (ICLR). Vienna, 2024. [doi: 10.48550/arXiv.2306.08543]
[10] Hinton G, Vinyals O, Dean J. Distilling the knowledge in a neural network. In: Proc. of the 2014 Deep Learning and Representation
Learning Workshop in Conjunction with NIPS. 2014. [doi: 10.48550/arXiv.1503.02531]
[11] Li TH, Li JG, Liu Z, Zhang CS. Few sample knowledge distillation for efficient network compression. In: Proc. of the 2020 IEEE/CVF
Conf. on Computer Vision and Pattern Recognition. Seattle: IEEE, 2020. 14627–14635. [doi: 10.1109/cvpr42600.2020.01465]
[12] Zagoruyko S, Komodakis N. Paying more attention to attention: Improving the performance of convolutional neural networks via
attention transfer. In: Proc. of the 5th Int’l Conf. on Learning Representations (ICLR). Toulon, 2017. [doi: 10.48550/arXiv.1612.03928]
[13] Passalis N, Tefas A. Learning deep representations with probabilistic knowledge transfer. In: Proc. of the 15th European Conf. on
Computer Vision. Munich: Springer, 2018. 283–299. [doi: 10.1007/978-3-030-01252-6_17]
[14] Huang ZH, Wang NY. Like what you like: Knowledge distill via neuron selectivity transfer. arXiv:1707.01219, 2017.
[15] Asif U, Tang JB, Harrer S. Ensemble knowledge distillation for learning improved and efficient networks. In: Proc. of the 2020 European
Conf. on Artificial Intelligence. IOS Press, 2020. 953–960. [doi: 10.3233/FAIA200188]
[16] Chen JY, Ren DD, Li WB, Huo J, Gao Y. Lightweight knowledge distillation for few-shot learning. Ruan Jian Xue Bao/Journal of
Software, 2024, 35(5): 2414–2429 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/6958.htm [doi: 10.13328/j.cnki.
jos.006958]
[17] Malinin A, Gales M. Reverse KL-divergence training of prior networks: Improved uncertainty and adversarial robustness. In: Proc. of the
33rd Int’l Conf. on Neural Information Processing Systems. Vancouver: Curran Associates Inc., 2019. 14547–14558.
[18] Allal LB, Li R, Kocetkov D, et al. SantaCoder: Don’t reach for the stars! arXiv:2301.03988, 2023.
[19] Rozière B, Gehring J, Gloeckle F, et al. Code Llama: Open foundation models for code. arXiv:2308.12950, 2024.
[20] Nijkamp E, Pang B, Hayashi H, Tu LF, Wang H, Zhou YB, Savarese S, Xiong CM. CodeGen: An open large language model for code
with multi-turn program synthesis. In: Proc. of the 11th Int’l Conf. on Learning Representations (ICLR). Kigali, 2023. [doi: 10.48550/
arXiv.2203.13474]
[21] Yuan ZQ, Lou YL, Liu MW, Ding SJ, Wang KX, Chen YX, Peng X. No more manual tests? Evaluating and improving ChatGPT for unit
test generation. arXiv:2305.04207, 2024.
[22] Chen ZM, Kommrusch S, Monperrus M. Neural transfer learning for repairing security vulnerabilities in C code. IEEE Trans. on
Software Engineering, 2023, 49(1): 147–165. [doi: 10.1109/TSE.2022.3147265]
[23] Harman M. The role of artificial intelligence in software engineering. In: Proc. of the 1st Int’l Workshop on Realizing AI Synergies in
Software Engineering (RAISE). Zurich: IEEE, 2012. 1–6. [doi: 10.1109/RAISE.2012.6227961]
[24] Zhang QJ, Zhang TK, Zhai J, Fang CR, Yu BW, Sun WS, Chen ZY. A critical review of large language model on software engineering:
An example from ChatGPT and automated program repair. arXiv:2310.08879, 2024.
[25] Li ZY, Lu S, Guo DY, Duan N, Jannu S, Jenks G, Majumder D, Green J, Svyatkovskiy A, Fu SY, Sundaresan N. Automating code
review activities by large-scale pre-training. In: Proc. of the 30th ACM Joint European Software Engineering Conf. and Symp. on the
Foundations of Software Engineering. Singapore: ACM, 2022. 1035–1047. [doi: 10.1145/3540250.3549081]
[26] Borgeaud S, Mensch A, Hoffmann J, et al. Improving language models by retrieving from trillions of tokens. In: Proc. of the 39th Int’l
Conf. on Machine Learning. Baltimore: PMLR, 2022. 2206–2240.
[27] Tao CY, Wu W, Xu C, Hu WP, Zhao DY, Yan R. One time of interaction may not be enough: Go deep with an interaction-over-
interaction network for response selection in dialogues. In: Proc. of the 57th Annual Meeting of the Association for Computational
Linguistics. Florence: ACL, 2019. 1–11. [doi: 10.18653/v1/P19-1001]
[28] Huang CS, Liu Q, Lin BY, Pang TY, Du C, Lin M. LoraHub: Efficient cross-task generalization via dynamic LoRA composition. In:
Proc. of the 1st Conf. on Language Modeling. 2024. [doi: 10.48550/arXiv.2307.13269]
[29] Li XL, Liang P. Prefix-tuning: Optimizing continuous prompts for generation. In: Proc. of the 59th Annual Meeting of the ACL and the
11th Int’l Joint Conf. on Natural Language Processing. 2021. [doi: 10.48550/arXiv.2101.00190]
[30] Lei T, Bai JW, Brahma S, Ainslie J, Lee K, Zhou YQ, Du N, Zhao VY, Wu YX, Li B, Zhang Y, Chang MW. Conditional adapters:
Parameter-efficient transfer learning with fast inference. In: Proc. of the 37th Int’l Conf. on Neural Information Processing Systems. New
Orleans: Curran Associates Inc., 2023. 8152–8172.
[31] Stiennon N, Ouyang L, Wu J, Ziegler DM, Lowe R, Voss C, Radford A, Amodei D, Christiano P. Learning to summarize from human
feedback. In: Proc. of the 34th Int’l Conf. on Neural Information Processing Systems. Vancouver: Curran Associates Inc., 2020.
3008–3021.

