Page 103 - 《软件学报》2026年第5期
P. 103

软件学报 ISSN 1000-9825, CODEN RUXUEW                                        E-mail: jos@iscas.ac.cn
                 2026,37(5):1982−2005 [doi: 10.13328/j.cnki.jos.007539] [CSTR: 32375.14.jos.007539]  http://www.jos.org.cn
                 ©中国科学院软件研究所版权所有.                                                          Tel: +86-10-62562563



                                                              *
                 姿态控制人物图像视频生成技术综述

                 李玘芮  1 ,    励雪巍  2 ,    赵    奇  1 ,    李    杰  1 ,    李    玺  1


                 1
                  (浙江大学 计算机科学与技术学院, 浙江 杭州 310013)
                 2
                  (上海电机学院 电子信息学院, 上海 201306)
                 通信作者: 李玺, E-mail: xilizju@zju.edu.cn

                 摘 要: 生成技术的飞速发展揭示了相关技术在实际应用中的潜力, 姿态控制人物图像视频生成技术                                (pose-guided
                 person image and video generation) 的核心目标是将输入信息的人物转换为指定姿态, 同时保持人物外观的高度一
                 致性. 该技术可以广泛应用于虚拟试穿与时尚行业、广告内容生成领域的视频生成与编辑以及多模态结合生成等
                 多个应用场景, 推动用户体验和技术创新的进步. 尽管该技术已经取得了显著进展, 仍面临着多个挑战, 包括姿态
                 迁移过程中外观信息的有效提取和重排、不可见信息的生成、一致性保持、模型的高效训练与使用等. 基于现有
                 技术的挑战, 详细分析了当前主流的姿态控制生成方法应对挑战的策略, 并探讨了它们在实际应用中的可行性和
                 局限性. 同时, 还讨论了姿态控制生成技术的常用生成模型以及不同的姿态信息表示方法. 此外, 整理讨论了该技
                 术常用的数据集大小、特点等信息、各项测试基准, 并从虚拟试穿、视频生成与编辑、多模态结合生成等应用场
                 景展开了讨论. 此外, 还揭示了目前方法仍遇到的个性化信息的保留、复杂场景的生成以及模型效率与实时性能
                 等挑战, 并讨论姿态控制生成技术可能的未来发展趋势, 旨在为相关领域的研究人员提供系统的总结与参考, 以期
                 推动该技术在各行业中的应用与创新.
                 关键词: 姿态控制生成; 人物图像视频生成; 生成对抗网络; 扩散模型; 可控生成
                 中图法分类号: TP181

                 中文引用格式: 李玘芮, 励雪巍, 赵奇, 李杰, 李玺. 姿态控制人物图像视频生成技术综述. 软件学报, 2026, 37(5): 1982–2005. http://
                 www.jos.org.cn/1000-9825/7539.htm
                 英文引用格式: Li QR, Li XW, Zhao Q, Li J, Li X. Review on Pose-guided Person Image and Video Generation Technologies. Ruan
                 Jian Xue Bao/Journal of Software, 2026, 37(5): 1982–2005 (in Chinese). http://www.jos.org.cn/1000-9825/7539.htm

                 Review on Pose-guided Person Image and Video Generation Technologies
                                                1
                                           1
                                  2
                        1
                 LI Qi-Rui , LI Xue-Wei , ZHAO Qi , LI Jie , LI Xi 1
                 1
                 (College of Computer Science and Technology, Zhejiang University, Hangzhou 310013, China)
                 2
                 (School of Electronic Information Engineering, Shanghai Dianji University, Shanghai 201306, China)
                 Abstract:  The  rapid  development  of  generative  technologies  has  revealed  their  potential  for  real-world  applications.  The  core  objective  of
                 pose-guided  person  image  and  video  generation  is  to  transform  a  person  from  inputs  into  a  specified  pose  while  maintaining  a  high  level
                 of  appearance  consistency.  This  technology  can  be  widely  applied  in  various  fields  such  as  virtual  try-on  and  fashion,  advertising  video
                 generation  and  editing,  and  multimodal  content  creation,  driving  advancements  in  user  experience  and  technological  innovation.  However,
                 despite  significant  progress,  the  technology  still  faces  multiple  challenges,  including  effective  extraction  and  rearrangement  of  appearance
                 information  during  pose  transfer,  generation  of  unseen  information,  consistency  preservation,  and  efficient  model  training  and  deployment.
                 Based  on  the  existing  challenges,  this  study  provides  a  detailed  analysis  of  the  strategies  employed  by  current  mainstream  pose-guided
                 generation  methods  to  address  these  issues,  discussing  their  feasibility  and  limitations  in  practical  applications.  Moreover,  it  explores  the


                 *    本文由“多媒体智能理解与生成”专题特约编辑孙立峰教授、闵巍庆副研究员、马占宇教授、蒋树强研究员、彭宇新教授、田丰研究
                  员、黄庆明教授推荐.
                  收稿时间: 2025-05-19; 修改时间: 2025-07-11; 采用时间: 2025-09-05; jos 在线出版时间: 2025-09-23
                  CNKI 网络首发时间: 2025-12-26
   98   99   100   101   102   103   104   105   106   107   108