Page 103 - 《软件学报》2026年第5期
P. 103
软件学报 ISSN 1000-9825, CODEN RUXUEW E-mail: jos@iscas.ac.cn
2026,37(5):1982−2005 [doi: 10.13328/j.cnki.jos.007539] [CSTR: 32375.14.jos.007539] http://www.jos.org.cn
©中国科学院软件研究所版权所有. Tel: +86-10-62562563
*
姿态控制人物图像视频生成技术综述
李玘芮 1 , 励雪巍 2 , 赵 奇 1 , 李 杰 1 , 李 玺 1
1
(浙江大学 计算机科学与技术学院, 浙江 杭州 310013)
2
(上海电机学院 电子信息学院, 上海 201306)
通信作者: 李玺, E-mail: xilizju@zju.edu.cn
摘 要: 生成技术的飞速发展揭示了相关技术在实际应用中的潜力, 姿态控制人物图像视频生成技术 (pose-guided
person image and video generation) 的核心目标是将输入信息的人物转换为指定姿态, 同时保持人物外观的高度一
致性. 该技术可以广泛应用于虚拟试穿与时尚行业、广告内容生成领域的视频生成与编辑以及多模态结合生成等
多个应用场景, 推动用户体验和技术创新的进步. 尽管该技术已经取得了显著进展, 仍面临着多个挑战, 包括姿态
迁移过程中外观信息的有效提取和重排、不可见信息的生成、一致性保持、模型的高效训练与使用等. 基于现有
技术的挑战, 详细分析了当前主流的姿态控制生成方法应对挑战的策略, 并探讨了它们在实际应用中的可行性和
局限性. 同时, 还讨论了姿态控制生成技术的常用生成模型以及不同的姿态信息表示方法. 此外, 整理讨论了该技
术常用的数据集大小、特点等信息、各项测试基准, 并从虚拟试穿、视频生成与编辑、多模态结合生成等应用场
景展开了讨论. 此外, 还揭示了目前方法仍遇到的个性化信息的保留、复杂场景的生成以及模型效率与实时性能
等挑战, 并讨论姿态控制生成技术可能的未来发展趋势, 旨在为相关领域的研究人员提供系统的总结与参考, 以期
推动该技术在各行业中的应用与创新.
关键词: 姿态控制生成; 人物图像视频生成; 生成对抗网络; 扩散模型; 可控生成
中图法分类号: TP181
中文引用格式: 李玘芮, 励雪巍, 赵奇, 李杰, 李玺. 姿态控制人物图像视频生成技术综述. 软件学报, 2026, 37(5): 1982–2005. http://
www.jos.org.cn/1000-9825/7539.htm
英文引用格式: Li QR, Li XW, Zhao Q, Li J, Li X. Review on Pose-guided Person Image and Video Generation Technologies. Ruan
Jian Xue Bao/Journal of Software, 2026, 37(5): 1982–2005 (in Chinese). http://www.jos.org.cn/1000-9825/7539.htm
Review on Pose-guided Person Image and Video Generation Technologies
1
1
2
1
LI Qi-Rui , LI Xue-Wei , ZHAO Qi , LI Jie , LI Xi 1
1
(College of Computer Science and Technology, Zhejiang University, Hangzhou 310013, China)
2
(School of Electronic Information Engineering, Shanghai Dianji University, Shanghai 201306, China)
Abstract: The rapid development of generative technologies has revealed their potential for real-world applications. The core objective of
pose-guided person image and video generation is to transform a person from inputs into a specified pose while maintaining a high level
of appearance consistency. This technology can be widely applied in various fields such as virtual try-on and fashion, advertising video
generation and editing, and multimodal content creation, driving advancements in user experience and technological innovation. However,
despite significant progress, the technology still faces multiple challenges, including effective extraction and rearrangement of appearance
information during pose transfer, generation of unseen information, consistency preservation, and efficient model training and deployment.
Based on the existing challenges, this study provides a detailed analysis of the strategies employed by current mainstream pose-guided
generation methods to address these issues, discussing their feasibility and limitations in practical applications. Moreover, it explores the
* 本文由“多媒体智能理解与生成”专题特约编辑孙立峰教授、闵巍庆副研究员、马占宇教授、蒋树强研究员、彭宇新教授、田丰研究
员、黄庆明教授推荐.
收稿时间: 2025-05-19; 修改时间: 2025-07-11; 采用时间: 2025-09-05; jos 在线出版时间: 2025-09-23
CNKI 网络首发时间: 2025-12-26

