Page 106 - 《软件学报》2026年第4期
P. 106
李昱洁 等: 面向开放世界持续学习的任务敏感提示驱动混合专家模型 1547
[30] Vaswani A, Bengio S, Brevdo E, Chollet F, Gomez A, Gouws S, Jones L, Kaiser Ł, Kalchbrenner N, Parmar N, Sepassi R, Shazeer N,
Uszkoreit J. Tensor2Tensor for neural machine translation. In: Proc. of the 13th Conf. of the Association for Machine Translation in the
Americas (Vol. 1: Research Track). Boston: Association for Machine Translation in the Americas, 2018. 193–199.
[31] Wang ZF, Zhang ZZ, Ebrahimi S, Sun RX, Zhang H, Lee CY, Ren XQ, Su GL, Perot V, Dy J, Pfister T. DualPrompt: Complementary
prompting for rehearsal-free continual learning. In: Proc. of the 17th European Conf. on Computer Vision. Tel Aviv: Springer, 2022.
631–648. [doi: 10.1007/978-3-031-19809-0_36]
[32] Wang YB, Huang ZW, Hong XP. S-prompts learning with pre-trained Transformers: An Occam’s razor for domain incremental learning.
In: Proc. of the 36th Int’l Conf. on Neural Information Processing Systems. New Orleans: Curran Associates Inc., 2022. 411.
[33] Smith JS, Karlinsky L, Gutta V, Cascante-Bonilla P, Kim D, Arbelle A, Panda R, Feris R, Kira Z. CODA-prompt: Continual decomposed
attention-based prompting for rehearsal-free continual learning. In: Proc. of the 2023 IEEE/CVF Conf. on Computer Vision and Pattern
Recognition. Vancouver: IEEE, 2023. 11909–11919. [doi: 10.1109/CVPR52729.2023.01146]
[34] Wang LY, Xie JY, Zhang XX, Huang MY, Su H, Zhu J. Hierarchical decomposition of prompt-based continual learning: Rethinking
obscured sub-optimality. In: Proc. of the 37th Int’l Conf. on Neural Information Processing Systems. New Orleans: Curran Associates
Inc., 2023. 3022.
[35] Shazeer N, Mirhoseini A, Maziarz K, Davis A, Le QV, Hinton GE, Dean J. Outrageously large neural networks: The sparsely-gated
mixture-of-experts layer. In: Proc. of the 5th Int’l Conf. on Learning Representations. Toulon: OpenReview.net, 2017. 1–19.
[36] Lepikhin D, Lee H, Xu YZ, Chen DH, Firat O, Huang YP, Krikun M, Shazeer N, Chen ZF. GShard: Scaling giant models with
conditional computation and automatic sharding. In: Proc. of the 9th Int’l Conf. on Learning Representations. Appleton: OpenReview.net,
2020. 1–23.
[37] Fedus W, Zoph B, Shazeer N. Switch Transformers: Scaling to trillion parameter models with simple and efficient sparsity. The Journal
of Machine Learning Research, 2022, 23(1): 120.
[38] Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I. Attention is all you need. In: Proc. of the
31st Int’l Conf. on Neural Information Processing Systems. Long Beach: Curran Associates Inc., 2017. 6000–6010.
[39] Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai XH, Unterthiner T, Dehghani M, Minderer M, Heigold G, Gelly S,
Uszkoreit J, Houlsby N. An image is worth 16x16 words: Transformers for image recognition at scale. In: Proc. of the 9th Int’l Conf. on
Learning Representations. Appleton: OpenReview.net, 2021. 1–21.
[40] McDonnell MD, Gong D, Parveneh A, Abbasnejad E, van den Hengel A. RanPAC: Random projections and pre-trained models for
continual learning. In: Proc. of the 37th Int’l Conf. on Neural Information Processing Systems. New Orleans: Curran Associates Inc.,
2023. 526.
[41] Zhou DW, Cai ZW, Ye HJ, Zhan DC, Liu ZW. Revisiting class-incremental learning with pre-trained models: Generalizability and
adaptivity are all you need. Int’l Journal of Computer Vision, 2025, 133(3): 1012–1032. [doi: 10.1007/s11263-024-02218-0]
[42] Hendrycks D, Basart S, Mazeika M, Zou A, Kwon J, Mostajabi M, Steinhardt J, Song D. Scaling out-of-distribution detection for real-
world settings. In: Proc. of the 39th Int’l Conf. on Machine Learning. Baltimore: PMLR, 2022. 8759–8773.
[43] Chan RB, Rottmann M, Gottschalk H. Entropy maximization and meta classification for out-of-distribution detection in semantic
segmentation. In: Proc. of the 2021 IEEE/CVF Int’l Conf. on Computer Vision. Montreal: IEEE, 2021. 5108–5117. [doi: 10.1109/
ICCV48922.2021.00508]
[44] Liu WT, Wang XY, Owens JD, Li YX. Energy-based out-of-distribution detection. In: Proc. of the 34th Int’l Conf. on Neural Information
Processing Systems. Vancouver: Curran Associates Inc., 2020. 1802.
作者简介
李昱洁, 博士生, CCF 学生会员, 主要研究领域为深度学习, 持续学习, 开放世界持续学习, 智能金融, 数据挖掘.
吴晗, 本科生, CCF 学生会员, 主要研究领域为深度学习, 持续学习, 开放世界持续学习.
孟丹, 博士, 教授, 博士生导师, 主要研究领域为智能金融, 智能决策, 不确定性信息处理, 数据挖掘.
李天瑞, 博士, 教授, 博士生导师, CCF 杰出会员, 主要研究领域为智能信息处理, 数据挖掘, 云计算, 大数据, 粗糙集与粒计算.
杨新, 博士, 教授, 博士生导师, CCF 杰出会员, 主要研究领域为机器学习, 联邦学习, 持续学习, 数据挖掘, 金融科技.

