Page 437 - 《软件学报》2026年第4期
P. 437

1878                                                       软件学报  2026  年第  37  卷第  4  期


                 素生成数据, 虽可控性强但效率较低; 而非自回归方法                (如扩散模型) 并行生成, 速度更快但可能牺牲一致性. 最新
                 研究如   Transfusion [123] 和  Show-O [124] 尝试结合两者优势, 使用自回归-扩散混合建模, 同时处理理解和生成任务. 为
                 适应任务的不同输入数据和变化, 它们采用文本编码器和图像编码器将文本和图像各自编码为离散                                 token, 并进一
                 步提出了统一的提示策略. 它们将这些            token  作为输入处理, 最终得到了图像和文本统一生成的效果. 未来, 生成
                 模型可能进一步融合不同范式, 在可控性、速度和生成能力上实现更优平衡.
                  6   总 结

                    本文系统综述了基于扩散模型的个性化图像生成方法, 通过相关文献的统计得知, 个性化生成技术已成为计
                 算机视觉领域的研究热点, 其核心目标是通过结合用户提供的特定概念与文本描述, 实现高质量、可控的定制化
                 图像生成. 本文首先从理论框架层面, 明确了扩散模型的数学基础与个性化生成的任务定义, 提出将生成过程标准
                 化为概念反演和个性化生成两阶段的统一范式. 其次, 针对当前研究从单主体向多概念发展的趋势, 本文分别对两
                 大方向展开深度分析: 1) 单主体生成方面, 本文介绍了物体驱动学习、人像驱动学习以及图像驱动个性化视频生
                 成这  3  个子任务, 现有方法通过设计概念反演和特征注入技术, 显著提升了主体视觉特征的保真度                           (如物体、人像
                 等细节重建); 2) 多概念生成任务方面, 本文从多概念生成与组合、多概念属性可控编辑和多概念服装组合试衣
                 这  3  方面进行综述, 这些任务聚焦于概念的语义对齐与视觉一致性, 通过分层控制、注意力机制优化、概念解耦
                 等手段解决了多主体融合的冲突问题. 本文进一步总结了常用数据集和评估指标, 并对比了不同方法的生成质量,
                 为后续研究提供了基准参考. 最后, 本文整理当前个性化生成方法仍存在的挑战, 并基于此给出了未来可能的研究
                 方向. 本文的综述为研究者提供了系统性的技术梳理与问题分析, 有望推动个性化生成技术在艺术创作、虚拟试
                 衣等场景的落地, 并促进生成模型领域的可持续发展.

                 References
                  [1]   Goodfellow IJ, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, Courville A, Bengio Y. Generative adversarial nets. In:
                      Proc. of the 28th Int’l Conf. on Neural Information Processing Systems. Montreal: MIT Press, 2014. 2672–2680.
                  [2]   Rombach R, Blattmann A, Lorenz D, Esser P, Ommer B. High-resolution image synthesis with latent diffusion models. In: Proc. of the
                      2022 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022. 10674–10685. [doi: 10.1109/CVPR
                      52688.2022.01042]
                  [3]   Zhu JY, Krähenbühl P, Shechtman E, Efros AA. Generative visual manipulation on the natural image manifold. In: Proc. of the 14th
                      European Conf. on Computer Vision. Amsterdam: Springer, 2016. 597–613. [doi: 10.1007/978-3-319-46454-1_36]
                  [4]   Richardson E, Alaluf Y, Patashnik O, Nitzan Y, Azar Y, Shapiro S, Cohen-Or D. Encoding in style: A StyleGAN encoder for image-to-
                      image  translation.  In:  Proc.  of  the  2021  IEEE/CVF  Conf.  on  Computer  Vision  and  Pattern  Recognition.  Nashville:  IEEE,  2021.
                      2287–2296. [doi: 10.1109/CVPR46437.2021.00232]
                  [5]   Tov  O,  Alaluf  Y,  Nitzan  Y,  Patashnik  O,  Cohen-Or  D.  Designing  an  encoder  for  StyleGAN  image  manipulation.  ACM  Trans.  on
                      Graphics, 2021, 40(4): 133. [doi: 10.1145/3450626.3459838]
                  [6]   Liu M, Wei YX, Wu XH, Zuo WM, Zhang L. Survey on leveraging pre-trained generative adversarial networks for image editing and
                      restoration. Science China Information Sciences, 2023, 66(5): 151101. [doi: 10.1007/s11432-022-3679-0]
                  [7]   Sohl-Dickstein J, Weiss EA, Maheswaranathan N, Ganguli S. Deep unsupervised learning using nonequilibrium thermodynamics. In:
                      Proc. of the 32nd Int’l Conf. on Machine Learning. Lille: PMLR, 2015. 2256–2265.
                  [8]   Ho J, Jain A, Abbeel P. Denoising diffusion probabilistic models. In: Proc. of the 34th Int’l Conf. on Neural Information Processing
                      Systems. Vancouver: Curran Associates Inc., 2020. 574.
                  [9]   Song  J,  Meng  CL,  Ermon  S.  Denoising  diffusion  implicit  models.  In:  Proc.  of  the  9th  Int’l  Conf.  on  Learning  Representations.
                      OpenReview.net, 2021. 1–20.
                 [10]   Zhang LM, Rao AY, Agrawala M. Adding conditional control to text-to-image diffusion models. In: Proc. of the 2023 IEEE/CVF Int’l
                      Conf. on Computer Vision. Paris: IEEE, 2023. 3813–3824. [doi: 10.1109/ICCV51070.2023.00355]
                 [11]   Zuo R, Hu HX, Deng XM, Ma CX, Wang HA. Survey on deep learning methods for freehand-sketch-based visual content generation.
                      Ruan Jian Xue Bao/Journal of Software, 2024, 35(7): 3497–3530 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/
                      7053.htm [doi: 10.13328/j.cnki.jos.007053]
   432   433   434   435   436   437   438   439   440   441   442