Page 442 - 《软件学报》2026年第4期
P. 442

何子健 等: 基于扩散模型的个性化图像生成方法综述                                                       1883


                 [94]   Heusel M, Ramsauer H, Unterthiner T, Nessler B, Hochreiter S. GANs trained by a two time-scale update rule converge to a local Nash
                      equilibrium.  In:  Proc.  of  the  31st  Conf.  on  Neural  Information  Processing  Systems.  Long  Beach:  Curran  Associates  Inc.,  2017.
                      6629–6640.
                 [95]   Szegedy C, Vanhoucke V, Ioffe S, Shlens J, Wojna Z. Rethinking the inception architecture for computer vision. In: Proc. of the 2016
                      IEEE Conf. on Computer Vision and Pattern Recognition. Las Vegas: IEEE, 2016. 2818–2826. [doi: 10.1109/CVPR.2016.308]
                 [96]   Unterthiner T, van Steenkiste S, Kurach K, Marinier R, Michalski M, Gelly S. Towards accurate generative models of video: A new
                      metric & challenges. arXiv:1812.01717, 2018.
                 [97]   Wang Z, Bovik AC, Sheikh HR, Simoncelli EP. Image quality assessment: From error visibility to structural similarity. IEEE Trans. on
                      Image Processing, 2004, 13(4): 600–612. [doi: 10.1109/TIP.2003.819861]
                 [98]   Zhang R, Isola P, Efros AA, Shechtman E, Wang O. The unreasonable effectiveness of deep features as a perceptual metric. In: Proc. of
                      the 2018 IEEE Conf. on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018. 586–595. [doi: 10.1109/CVPR.2018.
                      00068]
                 [99]   Kumari N, Zhang BL, Zhang R, Shechtman E, Zhu JY. Multi-concept customization of text-to-image diffusion. In: Proc. of the 2023
                      IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023. 1931–1941. [doi: 10.1109/CVPR52729.2023.
                      00192]
                 [100]   Zhang XL, Wei XY, Wu JL, Zhang TY, Zhang ZX, Lei Z, Li Q. Compositional inversion for stable diffusion models. In: Proc. of the
                      38th AAAI Conf. on Artificial Intelligence. Vancouver: AAAI Press, 2024. 7350–7358. [doi: 10.1609/aaai.v38i7.28565]
                 [101]   Pang LY, Yin J, Xie HR, Wang QP, Li Q, Mao XD. Cross initialization for personalized text-to-image generation. arXiv:2312.15905,
                      2023.
                 [102]   Hao SZ, Han K, Zhao SH, Wong KYK. ViCo: Plug-and-play visual condition for personalized text-to-image generation. arXiv:2306.
                      00971, 2023.
                 [103]   Zhang XL, Zhang WY, Wei XY, Wu JL, Zhang ZX, Lei Z, Li Q. Generative active learning for image synthesis personalization. In:
                      Proc. of the 32nd ACM Int’l Conf. on Multimedia. Melbourne: ACM, 2024. 10669–10677. [doi: 10.1145/3664647.3680773]
                 [104]   Tewel Y, Gal R, Chechik G, Atzmon Y. Key-locked rank one editing for text-to-image personalization. In: Proc. of the 2023 ACM
                      SIGGRAPH Conf. Los Angeles: ACM, 2023. 12. [doi: 10.1145/3588432.3591506]
                 [105]   Cao P, Yang L, Zhou F, Huang TR, Song Q. Concept-centric personalization with large-scale diffusion priors. arXiv:2312.08195, 2023.
                 [106]   Fang Y, Wang WJ, Zhang Y, Zhu FB, Wang QF, Feng FL, He XN. Reason4Rec: Large language models for recommendation with
                      deliberative user preference alignment. arXiv:2502.02061, 2025.
                 [107]   van  Le  T,  Phung  H,  Nguyen  TH,  Dao  Q,  Tran  NN,  Tran  A.  Anti-DreamBooth:  Protecting  users  from  personalized  text-to-image
                      synthesis. In: Proc. of the 2023 IEEE/CVF Int’l Conf. on Computer Vision. Paris: IEEE, 2023. 2116–2127. [doi: 10.1109/ICCV51070.
                      2023.00202]
                 [108]   Song YR, Yang P, Ci H, Shou MZ. IDProtector: An adversarial noise encoder to protect against ID-preserving image generation. In:
                      Proc. of the 2025 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Nashville: IEEE, 2025. 3019–3028. [doi: 10.1109/
                      CVPR52734.2025.00287]
                 [109]   Jones M, Wang SY, Kumari N, Bau D, Zhu JY. Customizing text-to-image models with a single image pair. In: Proc. of the 2024
                      SIGGRAPH Asia Conf. Papers. Tokyo: ACM, 2024. 6. [doi: 10.1145/3680528.3687642]
                 [110]   She D, Liu MS, Pang JX, Wang J, Yang Z, He WG, Zhang GH, Wang Y, Huang QH, Tang HB, Yu YL, Fu SM. CustomVideoX: 3D
                      reference attention driven dynamic adaptation for zero-shot customized video diffusion Transformers. arXiv:2502.06527, 2025.
                 [111]   Zhang YM, Xing ZN, Zeng YH, Fang YQ, Chen K. PIA: Your personalized image animator via plug-and-play modules in text-to-image
                      models. In: Proc. of the 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024. 7747–7756. [doi: 10.
                      1109/CVPR52733.2024.00740]
                 [112]   Liang F, Ma HY, He ZC, Hou TB, Hou J, Li KP, Dai XL, Juefei-Xu F, Azadi S, Sinha A, Zhang PZ, Vajda P, Marculescu D. Movie
                      Weaver: Tuning-free multi-concept video personalization with anchored prompts. In: Proc. of the 2025 IEEE/CVF Conf. on Computer
                      Vision and Pattern Recognition. Nashville: IEEE, 2025. 13146–13156. [doi: 10.1109/CVPR52734.2025.01227]
                 [113]   Li XM, Jia X, Wang QH, Diao HW, Ge MM, Li PX, He Y, Lu HC. MoTrans: Customized motion transfer with text-driven video
                      diffusion models. In: Proc. of the 32nd ACM Int’l Conf. on Multimedia. Melbourne: ACM, 2024. 3421–3430. [doi: 10.1145/3664647.
                      3680718]
                 [114]   Li HJ, Qiu HN, Zhang SW, Wang X, Wei YJ, Li ZK, Zhang YY, Wu BX, Cai D. PersonalVideo: High ID-fidelity video customization
                      without dynamic and semantic degradation. arXiv:2411.17048, 2024.
                 [115]   Wang  Z,  Li  AX,  Zhu  LT,  Guo  Y,  Dou  Q,  Li  ZG.  CustomVideo:  Customizing  text-to-video  generation  with  multiple  subjects.
   437   438   439   440   441   442   443   444