Page 438 - 《软件学报》2026年第4期
P. 438

何子健 等: 基于扩散模型的个性化图像生成方法综述                                                       1879


                 [12]   Yu K, Bin Y, Zheng ZQ, Yang Y. Text-to-image generation with conditional semantic augmentation. Ruan Jian Xue Bao/Journal of
                      Software, 2024, 35(5): 2150–2164 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/7024.htm [doi: 10.13328/j.cnki.
                      jos.007024]
                 [13]   Ruiz N, Li YZ, Jampani V, Pritch Y, Rubinstein M, Aberman K. DreamBooth: Fine tuning text-to-image diffusion models for subject-
                      driven  generation.  In:  Proc.  of  the  2023  IEEE/CVF  Conf.  on  Computer  Vision  and  Pattern  Recognition.  Vancouver:  IEEE,  2023.
                      22500–22510. [doi: 10.1109/CVPR52729.2023.02155]
                 [14]   Gal R, Alaluf Y, Atzmon Y, Patashnik O, Bermano AH, Chechik G, Cohen-Or D. An image is worth one word: Personalizing text-to-
                      image generation using textual inversion. In: Proc. of the 11th Int’l Conf. on Learning Representations. Kigali: OpenReview.net, 2023.
                      1–31.
                 [15]   Chen H, Zhang YP, Wu SM, Wang X, Duan XG, Zhou YW, Zhu WW. DisenBooth: Identity-preserving disentangled tuning for subject-
                      driven text-to-image generation. In: Proc. of the 12th Int’l Conf. on Learning Representations. Vienna: OpenReview.net, 2024. 1–23.
                 [16]   Radford  A,  Kim  JW,  Hallacy  C,  Ramesh  A,  Goh  G,  Agarwal  S,  Sastry  G,  Askell  A,  Mishkin  P,  Clark  J,  Krueger  G,  Sutskever  I.
                      Learning transferable visual models from natural language supervision. In: Proc. of the 38th Int’l Conf. on Machine Learning. Kigali:
                      PMLR, 2021. 8748–8763.
                 [17]   Li  JN,  Li  DX,  Xiong  CM,  Hoi  S.  BLIP:  Bootstrapping  language-image  pre-training  for  unified  vision-language  understanding  and
                      generation. In: Proc. of the 39th Int’l Conf. on Machine Learning. Baltimore: PMLR, 2022. 12888–12900.
                 [18]   Wei YX, Zhang YB, Ji ZL, Bai JF, Zhang L, Zuo WM. ELITE: Encoding visual concepts into textual embeddings for customized text-to-
                      image generation. In: Proc. of the 2023 IEEE/CVF Int’l Conf. on Computer Vision. Paris: IEEE, 2023. 15897–15907. [doi: 10.1109/
                      ICCV51070.2023.01461]
                 [19]   Li DX, Li JN, Hoi SCH. BLIP-Diffusion: Pre-trained subject representation for controllable text-to-image generation and editing. In:
                      Proc. of the 37th Int’l Conf. on Neural Information Processing Systems. New Orleans: Curran Associates Inc., 2023. 1312.
                 [20]   Ye H, Zhang J, Liu SB, Han X, Yang W. IP-Adapter: Text compatible image prompt adapter for text-to-image diffusion models. arXiv:
                      2308.06721, 2023.
                 [21]   Shi J, Xiong W, Lin Z, Jung HJ. InstantBooth: Personalized text-to-image generation without test-time finetuning. In: Proc. of the 2024
                      IEEE/CVF  Conf.  on  Computer  Vision  and  Pattern  Recognition.  Seattle:  IEEE,  2024.  8543–8552.  [doi: 10.1109/CVPR52733.2024.
                      00816]
                 [22]   Ma J, Liang JH, Chen C, Lu HN. Subject-Diffusion: Open domain personalized text-to-image generation without test-time fine-tuning.
                      In: Proc. of the 2024 ACM SIGGRAPH Conf. Denver: ACM, 2024. 25. [doi: 10.1145/3641519.3657469]
                 [23]   Jia XH, Zhao Y, Chan KCK, Li YD, Zhang H, Gong BQ, Hou TB, Wang HS, Su YC. Taming encoder for zero fine-tuning image
                      customization with text-to-image diffusion models. arXiv:2304.02642, 2023.
                 [24]   Xiao  ST,  Wang  YZ,  Zhou  JJ,  Yuan  HY,  Xing  XR,  Yan  RR,  Liu  Z.  OmniGen:  Unified  image  generation.  In:  Proc.  of  the  2025
                      IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Nashville: IEEE, 2025. 13294–13304. [doi: 10.1109/CVPR52734.2025.
                      01241]
                 [25]   Chen X, Zhang ZF, Zhang H, Zhou YQ, Kim SY, Liu Q, Li YJ, Zhang JM, Zhao NX, Wang YL, Ding H, Lin Z, Zhao HS. UniReal:
                      Universal image generation and editing via learning real-world dynamics. In: Proc. of the 2025 IEEE/CVF Conf. on Computer Vision
                      and Pattern Recognition. Nashville: IEEE, 2025. 12501–12511. [doi: 10.1109/CVPR52734.2025.01166]
                 [26]   Zhang  SL,  Huang  LH,  Chen  X,  Zhang  YF,  Wu  ZF,  Feng  YT,  Wang  W,  Shen  YJ,  Liu  Y,  Luo  P.  FlashFace:  Human  image
                      personalization with high-fidelity identity preservation. arXiv:2403.17008, 2024.
                 [27]   Li Z, Cao MD, Wang XT, Qi Z, Cheng MM, Shan Y. PhotoMaker: Customizing realistic human photos via stacked ID embedding. In:
                      Proc.  of  the  2024  IEEE/CVF  Conf.  on  Computer  Vision  and  Pattern  Recognition.  Seattle:  IEEE,  2024.  8640–8650.  [doi: 10.1109/
                      CVPR52733.2024.00825]
                 [28]   Deng  JK,  Guo  J,  Xue  NN,  Zafeiriou  S.  ArcFace:  Additive  angular  margin  loss  for  deep  face  recognition.  In:  Proc.  of  the  2019
                      IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2019. 4685–4694. [doi: 10.1109/CVPR.2019.00482]
                 [29]   Hu EJ, Shen YL, Wallis P, Allen-Zhu Z, Li YZ, Wang SA, Wang L, Chen WZ. LoRA: Low-rank adaptation of large language models.
                      In: Proc. of the 10th Int’l Conf. on Learning Representations. OpenReview.net, 2022. 1–13.
                 [30]   Wang QX, Bai X, Wang HF, Qin ZK, Chen A, Li HX, Tang X, Hu Y. InstantID: Zero-shot identity-preserving generation in seconds.
                      arXiv:2401.07519, 2024.
                 [31]   Peng X, Zhu JW, Jiang BY, Tai Y, Luo DH, Zhang J, Lin W, Jin TS, Wang CJ, Ji RR. PortraitBooth: A versatile portrait model for fast
                      identity-preserved personalization. In: Proc. of the 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Seattle: IEEE,
                      2024. 27070–27080. [doi: 10.1109/CVPR52733.2024.02557]
   433   434   435   436   437   438   439   440   441   442   443