Page 440 - 《软件学报》2026年第4期
P. 440
何子健 等: 基于扩散模型的个性化图像生成方法综述 1881
adaptions in frozen image-to-video diffusion model. In: Proc. of the 18th European Conf. on Computer Vision. Milan: Springer, 2024.
111–128. [doi: 10.1007/978-3-031-72655-2_7]
[54] Dai ZZ, Zhang ZH, Yao Y, Qiu BX, Zhu SY, Qin L, Wang WZ. AnimateAnything: Fine-grained open domain image animation with
motion guidance. arXiv:2311.12886, 2023.
[55] Avrahami O, Aberman K, Fried O, Cohen-Or D, Lischinski D. Break-A-Scene: Extracting multiple concepts from a single image. In:
Proc. of the 2024 SIGGRAPH Asia Conf. Sydney: ACM, 2023. 96. [doi: 10.1145/3610548.3618154]
[56] Vinker Y, Voynov A, Cohen-Or D, Shamir A. Concept decomposition for visual exploration and inspiration. ACM Trans. on Graphics,
2023, 42(6): 1–13. [doi: 10.1145/3618315]
[57] Hao SZ, Han K, Lv ZY, Zhao SH, Wong KYK. ConceptExpress: Harnessing diffusion models for single-image unsupervised concept
extraction. In: Proc. of the 18th European Conf. on Computer Vision. Milan: Springer, 2024. 215–233. [doi: 10.1007/978-3-031-73202-
7_13]
[58] Garibi D, Yadin S, Paiss R, Tov O, Zada S, Ephrat A, Michaeli T, Mosseri I, Dekel T. TokenVerse: Versatile multi-concept
personalization in token modulation space. ACM Trans. on Graphics, 2025, 44(4): 41.
[59] Peebles W, Xie SN. Scalable diffusion models with Transformers. In: Proc. of the 2023 IEEE/CVF Int’l Conf. on Computer Vision.
Paris: IEEE, 2023. 4172–4182. [doi: 10.1109/ICCV51070.2023.00387]
[60] Hu TH, Li LX, van de Weijer J, Gao HC, Shahbaz Khan F, Yang J, Cheng MM, Wang K, Wang YX. Token merging for training-free
semantic binding in text-to-image synthesis. In: Proc. of the 38th Int’l Conf. on Neural Information Processing Systems. Vancouver:
Curran Associates Inc., 2024. 4372.
[61] Yang Y, Wang W, Peng L, Song CT, Chen Y, Li HJ, Yang XL, Lu QL, Cai D, Wu BX, Liu W. LoRA-composer: Leveraging low-rank
adaptation for multi-concept customization in training-free diffusion models. arXiv:2403.11627, 2024.
[62] Kong Z, Zhang Y, Yang TY, Wang T, Zhang KH, Wu BZ, Chen GY, Liu W, Luo WH. OMG: Occlusion-friendly personalized multi-
concept generation in diffusion models. In: Proc. of the 18th European Conf. on Computer Vision. Milan: Springer, 2024. 253–270. [doi:
10.1007/978-3-031-72751-1_15]
[63] Gu YC, Wang XT, Wu JZ, Shi YJ, Chen YP, Fan ZH, Xiao WY, Zhao R, Chang SN, Wu WJ, Ge YX, Shan Y, Shou MZ. Mix-of-Show:
Decentralized low-rank adaptation for multi-concept customization of diffusion models. In: Proc. of the 37th Conf. on Neural
Information Processing Systems. New Orleans: Curran Associates Inc., 2023. 699.
[64] Po R, Yang GD, Aberman K, Wetzstein G. Orthogonal adaptation for modular customization of diffusion models. In: Proc. of the 2024
IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024. 7964–7973. [doi: 10.1109/CVPR52733.2024.
00761]
[65] Shah V, Ruiz N, Cole F, Lu E, Lazebnik S, Li YZ, Jampani V. ZipLoRA: Any subject in any style by effectively merging LoRAs. In:
Proc. of the 18th European Conf. on Computer Vision. Milan: Springer, 2024. 422–438. [doi: 10.1007/978-3-031-73232-4_24]
[66] Li YW, Li LG, Zhang ZY, Li XY, Wang GZ, Li HX, Cun XD, Shan Y, Zou YX. BlobCtrl: A unified and flexible framework for
element-level image generation and editing. arXiv:2503.13434, 2025.
[67] Chen X, Huang LH, Liu Y, Shen YJ, Zhao DL, Zhao HS. AnyDoor: Zero-shot object-level image customization. In: Proc. of the 2024
IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024. 6593–6602. [doi: 10.1109/CVPR52733.2024.
00630]
[68] Caron M, Touvron H, Misra I, Jégou H, Mairal J, Bojanowski P, Joulin A. Emerging properties in self-supervised vision Transformers.
In: Proc. of the 2021 IEEE/CVF Int’l Conf. on Computer Vision. Montreal: IEEE, 2021. 9630–9640. [doi: 10.1109/ICCV48922.2021.
00951]
[69] Mou C, Wang XT, Song JC, Shan Y, Zhang J. DragonDiffusion: Enabling drag-style manipulation on diffusion models. In: Proc. of the
12th Int’l Conf. on Learning Representations. Vienna: OpenReview.net, 2024. 1–12.
[70] Zhang ZY, Huang ZT, Liao J. Continuous layout editing of single images with diffusion models. Computer Graphics Forum, 2023,
42(7): e14966. [doi: 10.1111/cgf.14966]
[71] Li YH, Liu HT, Wu QY, Mu FZ, Yang JW, Gao JF, Li CY, Lee YJ. Gligen: Open-set grounded text-to-image generation. In: Proc. of
the 2023 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023. 22511–22521. [doi: 10.1109/CVPR
52729.2023.02156]
[72] Alzayer H, Xia ZH, Zhang X, Shechtman E, Huang JB, Gharbi M. Magic Fixup: Streamlining photo editing by watching dynamic
videos. arXiv:2403.13044, 2024.
[73] Wang XR, Fu SM, Huang QH, He WG, Jiang H. MS-Diffusion: Multi-subject zero-shot image personalization with layout guidance. In:
Proc. of the 13th Int’l Conf. on Learning Representations. Singapore: OpenReview.net, 2025. 1–29.

