Page 438 - 《软件学报》2026年第4期
P. 438
何子健 等: 基于扩散模型的个性化图像生成方法综述 1879
[12] Yu K, Bin Y, Zheng ZQ, Yang Y. Text-to-image generation with conditional semantic augmentation. Ruan Jian Xue Bao/Journal of
Software, 2024, 35(5): 2150–2164 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/7024.htm [doi: 10.13328/j.cnki.
jos.007024]
[13] Ruiz N, Li YZ, Jampani V, Pritch Y, Rubinstein M, Aberman K. DreamBooth: Fine tuning text-to-image diffusion models for subject-
driven generation. In: Proc. of the 2023 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023.
22500–22510. [doi: 10.1109/CVPR52729.2023.02155]
[14] Gal R, Alaluf Y, Atzmon Y, Patashnik O, Bermano AH, Chechik G, Cohen-Or D. An image is worth one word: Personalizing text-to-
image generation using textual inversion. In: Proc. of the 11th Int’l Conf. on Learning Representations. Kigali: OpenReview.net, 2023.
1–31.
[15] Chen H, Zhang YP, Wu SM, Wang X, Duan XG, Zhou YW, Zhu WW. DisenBooth: Identity-preserving disentangled tuning for subject-
driven text-to-image generation. In: Proc. of the 12th Int’l Conf. on Learning Representations. Vienna: OpenReview.net, 2024. 1–23.
[16] Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, Krueger G, Sutskever I.
Learning transferable visual models from natural language supervision. In: Proc. of the 38th Int’l Conf. on Machine Learning. Kigali:
PMLR, 2021. 8748–8763.
[17] Li JN, Li DX, Xiong CM, Hoi S. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and
generation. In: Proc. of the 39th Int’l Conf. on Machine Learning. Baltimore: PMLR, 2022. 12888–12900.
[18] Wei YX, Zhang YB, Ji ZL, Bai JF, Zhang L, Zuo WM. ELITE: Encoding visual concepts into textual embeddings for customized text-to-
image generation. In: Proc. of the 2023 IEEE/CVF Int’l Conf. on Computer Vision. Paris: IEEE, 2023. 15897–15907. [doi: 10.1109/
ICCV51070.2023.01461]
[19] Li DX, Li JN, Hoi SCH. BLIP-Diffusion: Pre-trained subject representation for controllable text-to-image generation and editing. In:
Proc. of the 37th Int’l Conf. on Neural Information Processing Systems. New Orleans: Curran Associates Inc., 2023. 1312.
[20] Ye H, Zhang J, Liu SB, Han X, Yang W. IP-Adapter: Text compatible image prompt adapter for text-to-image diffusion models. arXiv:
2308.06721, 2023.
[21] Shi J, Xiong W, Lin Z, Jung HJ. InstantBooth: Personalized text-to-image generation without test-time finetuning. In: Proc. of the 2024
IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024. 8543–8552. [doi: 10.1109/CVPR52733.2024.
00816]
[22] Ma J, Liang JH, Chen C, Lu HN. Subject-Diffusion: Open domain personalized text-to-image generation without test-time fine-tuning.
In: Proc. of the 2024 ACM SIGGRAPH Conf. Denver: ACM, 2024. 25. [doi: 10.1145/3641519.3657469]
[23] Jia XH, Zhao Y, Chan KCK, Li YD, Zhang H, Gong BQ, Hou TB, Wang HS, Su YC. Taming encoder for zero fine-tuning image
customization with text-to-image diffusion models. arXiv:2304.02642, 2023.
[24] Xiao ST, Wang YZ, Zhou JJ, Yuan HY, Xing XR, Yan RR, Liu Z. OmniGen: Unified image generation. In: Proc. of the 2025
IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Nashville: IEEE, 2025. 13294–13304. [doi: 10.1109/CVPR52734.2025.
01241]
[25] Chen X, Zhang ZF, Zhang H, Zhou YQ, Kim SY, Liu Q, Li YJ, Zhang JM, Zhao NX, Wang YL, Ding H, Lin Z, Zhao HS. UniReal:
Universal image generation and editing via learning real-world dynamics. In: Proc. of the 2025 IEEE/CVF Conf. on Computer Vision
and Pattern Recognition. Nashville: IEEE, 2025. 12501–12511. [doi: 10.1109/CVPR52734.2025.01166]
[26] Zhang SL, Huang LH, Chen X, Zhang YF, Wu ZF, Feng YT, Wang W, Shen YJ, Liu Y, Luo P. FlashFace: Human image
personalization with high-fidelity identity preservation. arXiv:2403.17008, 2024.
[27] Li Z, Cao MD, Wang XT, Qi Z, Cheng MM, Shan Y. PhotoMaker: Customizing realistic human photos via stacked ID embedding. In:
Proc. of the 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024. 8640–8650. [doi: 10.1109/
CVPR52733.2024.00825]
[28] Deng JK, Guo J, Xue NN, Zafeiriou S. ArcFace: Additive angular margin loss for deep face recognition. In: Proc. of the 2019
IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2019. 4685–4694. [doi: 10.1109/CVPR.2019.00482]
[29] Hu EJ, Shen YL, Wallis P, Allen-Zhu Z, Li YZ, Wang SA, Wang L, Chen WZ. LoRA: Low-rank adaptation of large language models.
In: Proc. of the 10th Int’l Conf. on Learning Representations. OpenReview.net, 2022. 1–13.
[30] Wang QX, Bai X, Wang HF, Qin ZK, Chen A, Li HX, Tang X, Hu Y. InstantID: Zero-shot identity-preserving generation in seconds.
arXiv:2401.07519, 2024.
[31] Peng X, Zhu JW, Jiang BY, Tai Y, Luo DH, Zhang J, Lin W, Jin TS, Wang CJ, Ji RR. PortraitBooth: A versatile portrait model for fast
identity-preserved personalization. In: Proc. of the 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Seattle: IEEE,
2024. 27070–27080. [doi: 10.1109/CVPR52733.2024.02557]

