Page 442 - 《软件学报》2026年第4期
P. 442
何子健 等: 基于扩散模型的个性化图像生成方法综述 1883
[94] Heusel M, Ramsauer H, Unterthiner T, Nessler B, Hochreiter S. GANs trained by a two time-scale update rule converge to a local Nash
equilibrium. In: Proc. of the 31st Conf. on Neural Information Processing Systems. Long Beach: Curran Associates Inc., 2017.
6629–6640.
[95] Szegedy C, Vanhoucke V, Ioffe S, Shlens J, Wojna Z. Rethinking the inception architecture for computer vision. In: Proc. of the 2016
IEEE Conf. on Computer Vision and Pattern Recognition. Las Vegas: IEEE, 2016. 2818–2826. [doi: 10.1109/CVPR.2016.308]
[96] Unterthiner T, van Steenkiste S, Kurach K, Marinier R, Michalski M, Gelly S. Towards accurate generative models of video: A new
metric & challenges. arXiv:1812.01717, 2018.
[97] Wang Z, Bovik AC, Sheikh HR, Simoncelli EP. Image quality assessment: From error visibility to structural similarity. IEEE Trans. on
Image Processing, 2004, 13(4): 600–612. [doi: 10.1109/TIP.2003.819861]
[98] Zhang R, Isola P, Efros AA, Shechtman E, Wang O. The unreasonable effectiveness of deep features as a perceptual metric. In: Proc. of
the 2018 IEEE Conf. on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018. 586–595. [doi: 10.1109/CVPR.2018.
00068]
[99] Kumari N, Zhang BL, Zhang R, Shechtman E, Zhu JY. Multi-concept customization of text-to-image diffusion. In: Proc. of the 2023
IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023. 1931–1941. [doi: 10.1109/CVPR52729.2023.
00192]
[100] Zhang XL, Wei XY, Wu JL, Zhang TY, Zhang ZX, Lei Z, Li Q. Compositional inversion for stable diffusion models. In: Proc. of the
38th AAAI Conf. on Artificial Intelligence. Vancouver: AAAI Press, 2024. 7350–7358. [doi: 10.1609/aaai.v38i7.28565]
[101] Pang LY, Yin J, Xie HR, Wang QP, Li Q, Mao XD. Cross initialization for personalized text-to-image generation. arXiv:2312.15905,
2023.
[102] Hao SZ, Han K, Zhao SH, Wong KYK. ViCo: Plug-and-play visual condition for personalized text-to-image generation. arXiv:2306.
00971, 2023.
[103] Zhang XL, Zhang WY, Wei XY, Wu JL, Zhang ZX, Lei Z, Li Q. Generative active learning for image synthesis personalization. In:
Proc. of the 32nd ACM Int’l Conf. on Multimedia. Melbourne: ACM, 2024. 10669–10677. [doi: 10.1145/3664647.3680773]
[104] Tewel Y, Gal R, Chechik G, Atzmon Y. Key-locked rank one editing for text-to-image personalization. In: Proc. of the 2023 ACM
SIGGRAPH Conf. Los Angeles: ACM, 2023. 12. [doi: 10.1145/3588432.3591506]
[105] Cao P, Yang L, Zhou F, Huang TR, Song Q. Concept-centric personalization with large-scale diffusion priors. arXiv:2312.08195, 2023.
[106] Fang Y, Wang WJ, Zhang Y, Zhu FB, Wang QF, Feng FL, He XN. Reason4Rec: Large language models for recommendation with
deliberative user preference alignment. arXiv:2502.02061, 2025.
[107] van Le T, Phung H, Nguyen TH, Dao Q, Tran NN, Tran A. Anti-DreamBooth: Protecting users from personalized text-to-image
synthesis. In: Proc. of the 2023 IEEE/CVF Int’l Conf. on Computer Vision. Paris: IEEE, 2023. 2116–2127. [doi: 10.1109/ICCV51070.
2023.00202]
[108] Song YR, Yang P, Ci H, Shou MZ. IDProtector: An adversarial noise encoder to protect against ID-preserving image generation. In:
Proc. of the 2025 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Nashville: IEEE, 2025. 3019–3028. [doi: 10.1109/
CVPR52734.2025.00287]
[109] Jones M, Wang SY, Kumari N, Bau D, Zhu JY. Customizing text-to-image models with a single image pair. In: Proc. of the 2024
SIGGRAPH Asia Conf. Papers. Tokyo: ACM, 2024. 6. [doi: 10.1145/3680528.3687642]
[110] She D, Liu MS, Pang JX, Wang J, Yang Z, He WG, Zhang GH, Wang Y, Huang QH, Tang HB, Yu YL, Fu SM. CustomVideoX: 3D
reference attention driven dynamic adaptation for zero-shot customized video diffusion Transformers. arXiv:2502.06527, 2025.
[111] Zhang YM, Xing ZN, Zeng YH, Fang YQ, Chen K. PIA: Your personalized image animator via plug-and-play modules in text-to-image
models. In: Proc. of the 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024. 7747–7756. [doi: 10.
1109/CVPR52733.2024.00740]
[112] Liang F, Ma HY, He ZC, Hou TB, Hou J, Li KP, Dai XL, Juefei-Xu F, Azadi S, Sinha A, Zhang PZ, Vajda P, Marculescu D. Movie
Weaver: Tuning-free multi-concept video personalization with anchored prompts. In: Proc. of the 2025 IEEE/CVF Conf. on Computer
Vision and Pattern Recognition. Nashville: IEEE, 2025. 13146–13156. [doi: 10.1109/CVPR52734.2025.01227]
[113] Li XM, Jia X, Wang QH, Diao HW, Ge MM, Li PX, He Y, Lu HC. MoTrans: Customized motion transfer with text-driven video
diffusion models. In: Proc. of the 32nd ACM Int’l Conf. on Multimedia. Melbourne: ACM, 2024. 3421–3430. [doi: 10.1145/3664647.
3680718]
[114] Li HJ, Qiu HN, Zhang SW, Wang X, Wei YJ, Li ZK, Zhang YY, Wu BX, Cai D. PersonalVideo: High ID-fidelity video customization
without dynamic and semantic degradation. arXiv:2411.17048, 2024.
[115] Wang Z, Li AX, Zhu LT, Guo Y, Dou Q, Li ZG. CustomVideo: Customizing text-to-video generation with multiple subjects.

