Page 106 - 《软件学报》2026年第4期
P. 106

李昱洁 等: 面向开放世界持续学习的任务敏感提示驱动混合专家模型                                                1547


                 [30]   Vaswani A, Bengio S, Brevdo E, Chollet F, Gomez A, Gouws S, Jones L, Kaiser Ł, Kalchbrenner N, Parmar N, Sepassi R, Shazeer N,
                     Uszkoreit J. Tensor2Tensor for neural machine translation. In: Proc. of the 13th Conf. of the Association for Machine Translation in the
                     Americas (Vol. 1: Research Track). Boston: Association for Machine Translation in the Americas, 2018. 193–199.
                 [31]   Wang ZF, Zhang ZZ, Ebrahimi S, Sun RX, Zhang H, Lee CY, Ren XQ, Su GL, Perot V, Dy J, Pfister T. DualPrompt: Complementary
                     prompting for rehearsal-free continual learning. In: Proc. of the 17th European Conf. on Computer Vision. Tel Aviv: Springer, 2022.
                     631–648. [doi: 10.1007/978-3-031-19809-0_36]
                 [32]   Wang YB, Huang ZW, Hong XP. S-prompts learning with pre-trained Transformers: An Occam’s razor for domain incremental learning.
                     In: Proc. of the 36th Int’l Conf. on Neural Information Processing Systems. New Orleans: Curran Associates Inc., 2022. 411.
                 [33]   Smith JS, Karlinsky L, Gutta V, Cascante-Bonilla P, Kim D, Arbelle A, Panda R, Feris R, Kira Z. CODA-prompt: Continual decomposed
                     attention-based prompting for rehearsal-free continual learning. In: Proc. of the 2023 IEEE/CVF Conf. on Computer Vision and Pattern
                     Recognition. Vancouver: IEEE, 2023. 11909–11919. [doi: 10.1109/CVPR52729.2023.01146]
                 [34]   Wang LY, Xie JY, Zhang XX, Huang MY, Su H, Zhu J. Hierarchical decomposition of prompt-based continual learning: Rethinking
                     obscured sub-optimality. In: Proc. of the 37th Int’l Conf. on Neural Information Processing Systems. New Orleans: Curran Associates
                     Inc., 2023. 3022.
                 [35]   Shazeer N, Mirhoseini A, Maziarz K, Davis A, Le QV, Hinton GE, Dean J. Outrageously large neural networks: The sparsely-gated
                     mixture-of-experts layer. In: Proc. of the 5th Int’l Conf. on Learning Representations. Toulon: OpenReview.net, 2017. 1–19.
                 [36]   Lepikhin  D,  Lee  H,  Xu  YZ,  Chen  DH,  Firat  O,  Huang  YP,  Krikun  M,  Shazeer  N,  Chen  ZF.  GShard:  Scaling  giant  models  with
                     conditional computation and automatic sharding. In: Proc. of the 9th Int’l Conf. on Learning Representations. Appleton: OpenReview.net,
                     2020. 1–23.
                 [37]   Fedus W, Zoph B, Shazeer N. Switch Transformers: Scaling to trillion parameter models with simple and efficient sparsity. The Journal
                     of Machine Learning Research, 2022, 23(1): 120.
                 [38]   Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I. Attention is all you need. In: Proc. of the
                     31st Int’l Conf. on Neural Information Processing Systems. Long Beach: Curran Associates Inc., 2017. 6000–6010.
                 [39]   Dosovitskiy  A,  Beyer  L,  Kolesnikov  A,  Weissenborn  D,  Zhai  XH,  Unterthiner  T,  Dehghani  M,  Minderer  M,  Heigold  G,  Gelly  S,
                     Uszkoreit J, Houlsby N. An image is worth 16x16 words: Transformers for image recognition at scale. In: Proc. of the 9th Int’l Conf. on
                     Learning Representations. Appleton: OpenReview.net, 2021. 1–21.
                 [40]   McDonnell  MD,  Gong  D,  Parveneh  A,  Abbasnejad  E,  van  den  Hengel  A.  RanPAC:  Random  projections  and  pre-trained  models  for
                     continual learning. In: Proc. of the 37th Int’l Conf. on Neural Information Processing Systems. New Orleans: Curran Associates Inc.,
                     2023. 526.
                 [41]   Zhou  DW,  Cai  ZW,  Ye  HJ,  Zhan  DC,  Liu  ZW.  Revisiting  class-incremental  learning  with  pre-trained  models:  Generalizability  and
                     adaptivity are all you need. Int’l Journal of Computer Vision, 2025, 133(3): 1012–1032. [doi: 10.1007/s11263-024-02218-0]
                 [42]   Hendrycks D, Basart S, Mazeika M, Zou A, Kwon J, Mostajabi M, Steinhardt J, Song D. Scaling out-of-distribution detection for real-
                     world settings. In: Proc. of the 39th Int’l Conf. on Machine Learning. Baltimore: PMLR, 2022. 8759–8773.
                 [43]   Chan  RB,  Rottmann  M,  Gottschalk  H.  Entropy  maximization  and  meta  classification  for  out-of-distribution  detection  in  semantic
                     segmentation.  In:  Proc.  of  the  2021  IEEE/CVF  Int’l  Conf.  on  Computer  Vision.  Montreal:  IEEE,  2021.  5108–5117.  [doi:  10.1109/
                     ICCV48922.2021.00508]
                 [44]   Liu WT, Wang XY, Owens JD, Li YX. Energy-based out-of-distribution detection. In: Proc. of the 34th Int’l Conf. on Neural Information
                     Processing Systems. Vancouver: Curran Associates Inc., 2020. 1802.

                 作者简介
                 李昱洁, 博士生, CCF  学生会员, 主要研究领域为深度学习, 持续学习, 开放世界持续学习, 智能金融, 数据挖掘.
                 吴晗, 本科生, CCF  学生会员, 主要研究领域为深度学习, 持续学习, 开放世界持续学习.
                 孟丹, 博士, 教授, 博士生导师, 主要研究领域为智能金融, 智能决策, 不确定性信息处理, 数据挖掘.
                 李天瑞, 博士, 教授, 博士生导师, CCF  杰出会员, 主要研究领域为智能信息处理, 数据挖掘, 云计算, 大数据, 粗糙集与粒计算.
                 杨新, 博士, 教授, 博士生导师, CCF  杰出会员, 主要研究领域为机器学习, 联邦学习, 持续学习, 数据挖掘, 金融科技.
   101   102   103   104   105   106   107   108   109   110   111