Page 437 - 《软件学报》2026年第4期
P. 437
1878 软件学报 2026 年第 37 卷第 4 期
素生成数据, 虽可控性强但效率较低; 而非自回归方法 (如扩散模型) 并行生成, 速度更快但可能牺牲一致性. 最新
研究如 Transfusion [123] 和 Show-O [124] 尝试结合两者优势, 使用自回归-扩散混合建模, 同时处理理解和生成任务. 为
适应任务的不同输入数据和变化, 它们采用文本编码器和图像编码器将文本和图像各自编码为离散 token, 并进一
步提出了统一的提示策略. 它们将这些 token 作为输入处理, 最终得到了图像和文本统一生成的效果. 未来, 生成
模型可能进一步融合不同范式, 在可控性、速度和生成能力上实现更优平衡.
6 总 结
本文系统综述了基于扩散模型的个性化图像生成方法, 通过相关文献的统计得知, 个性化生成技术已成为计
算机视觉领域的研究热点, 其核心目标是通过结合用户提供的特定概念与文本描述, 实现高质量、可控的定制化
图像生成. 本文首先从理论框架层面, 明确了扩散模型的数学基础与个性化生成的任务定义, 提出将生成过程标准
化为概念反演和个性化生成两阶段的统一范式. 其次, 针对当前研究从单主体向多概念发展的趋势, 本文分别对两
大方向展开深度分析: 1) 单主体生成方面, 本文介绍了物体驱动学习、人像驱动学习以及图像驱动个性化视频生
成这 3 个子任务, 现有方法通过设计概念反演和特征注入技术, 显著提升了主体视觉特征的保真度 (如物体、人像
等细节重建); 2) 多概念生成任务方面, 本文从多概念生成与组合、多概念属性可控编辑和多概念服装组合试衣
这 3 方面进行综述, 这些任务聚焦于概念的语义对齐与视觉一致性, 通过分层控制、注意力机制优化、概念解耦
等手段解决了多主体融合的冲突问题. 本文进一步总结了常用数据集和评估指标, 并对比了不同方法的生成质量,
为后续研究提供了基准参考. 最后, 本文整理当前个性化生成方法仍存在的挑战, 并基于此给出了未来可能的研究
方向. 本文的综述为研究者提供了系统性的技术梳理与问题分析, 有望推动个性化生成技术在艺术创作、虚拟试
衣等场景的落地, 并促进生成模型领域的可持续发展.
References
[1] Goodfellow IJ, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, Courville A, Bengio Y. Generative adversarial nets. In:
Proc. of the 28th Int’l Conf. on Neural Information Processing Systems. Montreal: MIT Press, 2014. 2672–2680.
[2] Rombach R, Blattmann A, Lorenz D, Esser P, Ommer B. High-resolution image synthesis with latent diffusion models. In: Proc. of the
2022 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022. 10674–10685. [doi: 10.1109/CVPR
52688.2022.01042]
[3] Zhu JY, Krähenbühl P, Shechtman E, Efros AA. Generative visual manipulation on the natural image manifold. In: Proc. of the 14th
European Conf. on Computer Vision. Amsterdam: Springer, 2016. 597–613. [doi: 10.1007/978-3-319-46454-1_36]
[4] Richardson E, Alaluf Y, Patashnik O, Nitzan Y, Azar Y, Shapiro S, Cohen-Or D. Encoding in style: A StyleGAN encoder for image-to-
image translation. In: Proc. of the 2021 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Nashville: IEEE, 2021.
2287–2296. [doi: 10.1109/CVPR46437.2021.00232]
[5] Tov O, Alaluf Y, Nitzan Y, Patashnik O, Cohen-Or D. Designing an encoder for StyleGAN image manipulation. ACM Trans. on
Graphics, 2021, 40(4): 133. [doi: 10.1145/3450626.3459838]
[6] Liu M, Wei YX, Wu XH, Zuo WM, Zhang L. Survey on leveraging pre-trained generative adversarial networks for image editing and
restoration. Science China Information Sciences, 2023, 66(5): 151101. [doi: 10.1007/s11432-022-3679-0]
[7] Sohl-Dickstein J, Weiss EA, Maheswaranathan N, Ganguli S. Deep unsupervised learning using nonequilibrium thermodynamics. In:
Proc. of the 32nd Int’l Conf. on Machine Learning. Lille: PMLR, 2015. 2256–2265.
[8] Ho J, Jain A, Abbeel P. Denoising diffusion probabilistic models. In: Proc. of the 34th Int’l Conf. on Neural Information Processing
Systems. Vancouver: Curran Associates Inc., 2020. 574.
[9] Song J, Meng CL, Ermon S. Denoising diffusion implicit models. In: Proc. of the 9th Int’l Conf. on Learning Representations.
OpenReview.net, 2021. 1–20.
[10] Zhang LM, Rao AY, Agrawala M. Adding conditional control to text-to-image diffusion models. In: Proc. of the 2023 IEEE/CVF Int’l
Conf. on Computer Vision. Paris: IEEE, 2023. 3813–3824. [doi: 10.1109/ICCV51070.2023.00355]
[11] Zuo R, Hu HX, Deng XM, Ma CX, Wang HA. Survey on deep learning methods for freehand-sketch-based visual content generation.
Ruan Jian Xue Bao/Journal of Software, 2024, 35(7): 3497–3530 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/
7053.htm [doi: 10.13328/j.cnki.jos.007053]

