Page 120 - 《软件学报》2026年第5期
P. 120
李玘芮 等: 姿态控制人物图像视频生成技术综述 1999
生成技术或可扩展至虚拟社交和个性化数字人生成领域. 例如, 结合表情控制和动态姿态, 虚拟社交平台中高度个
性化的虚拟人物形象能够为用户提供更加沉浸式的体验. 此外, 在社交媒体内容创作中, 生成技术有望进一步降低
内容创作者的技术门槛, 例如用于自动化生成个性化短视频或动态表情包.
6 总结和结论
姿态控制人物图像视频生成技术作为生成建模与计算机视觉领域的重要方向, 在虚拟试穿、广告创作、影视
制作等领域展现了巨大的应用潜力. 本文介绍了完成相关任务的主要生成器与常见姿态数据的表示方法, 并以解
决任务核心挑战为导向综述了相关技术的发展现状, 包括信息保留与重拍、推理不可见外观信息、一致性问题与
模型的高效训练和使用, 整理了相关数据集与任务评价指标, 同时分析了个性化信息保留、复杂背景生成以及提
升模型的实时性与效率等技术挑战并分析了其发展趋势.
References
[1] Wang T, Li LJ, Lin K, Zhai YH, Lin CC, Yang ZY, Zhang HW, Liu ZC, Wang LJ. DisCo: Disentangled control for realistic human
dance generation. In: Proc. of the 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). Seattle: IEEE, 2024.
9326–9336. [doi: 10.1109/CVPR52733.2024.00891]
[2] Yoon JS, Liu LJ, Golyanik V, Sarkar K, Park HS, Theobalt C. Pose-guided human animation from a single image in the wild. In: Proc.
of the 2021 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). Nashville: IEEE, 2021. 15034–15043. [doi: 10.
1109/cvpr46437.2021.01479]
[3] Wang X, Zhang SW, Gao CX, Wang JY, Zhou XQ, Zhang YY, Yan LX, Sang N. UniAnimate: Taming unified video diffusion models
for consistent human image animation. arXiv:2406.01188, 2024.
[4] Zhang Y, Gu JX, Wang LW, Wang H, Cheng JQ, Zhu YF, Zou FY. MimicMotion: High-quality human motion video generation with
confidence-aware pose guidance. arXiv:2406.19680, 2025.
[5] Dong HY, Liang XD, Gong K, Lai HJ, Zhu J, Yin J. Soft-gated warping-GAN for pose-guided person image synthesis. In: Proc. of the
32nd Int’l Conf. on Neural Information Processing Systems. Montréal: Curran Associates Inc., 2018. 472–482.
[6] Cui AY, McKee D, Lazebnik S. Dressing in order: Recurrent person image generation for pose transfer, virtual try-on and outfit editing.
In: Proc. of the 2021 IEEE/CVF Int’l Conf. on Computer Vision (ICCV). Montreal: IEEE, 2021. 14618–14627. [doi: 10.1109/iccv48922.
2021.01437]
[7] Sun K, Cao J, Wang Q, Tian LR, Zhang XD, Zhuo L, Zhang B, Bo LF, Zhou WB, Zhang WM, Gao DH. OutfitAnyone: Ultra-high
quality virtual try-on for any clothing and any person. arXiv:2407.16224, 2024.
[8] Fang NY, Qiu LM, Zhang SY, Wang ZL, Hu KR, Dong LY. A novel human image sequence synthesis method by pose-shape-content
inference. IEEE Trans. on Multimedia, 2023, 25: 6512–6524. [doi: 10.1109/TMM.2022.3209924]
[9] Shi CC, Chen YX, Lei BR, Chen JC. FashionPose: Text to pose to relight image generation for personalized fashion visualization.
arXiv:2507.13311, 2025.
[10] Randhavane T, Bera A, Kapsaskis K, Gray K, Manocha D. FVA: Modeling perceived friendliness of virtual agents using movement
characteristics. IEEE Trans. on Visualization and Computer Graphics, 2019, 25(11): 3135–3145. [doi: 10.1109/TVCG.2019.2932235]
[11] Xiang J, Guo YD, Hu LP, Guo BY, Yuan YC, Zhang JY. One shot, one talk: Whole-body talking avatar from a single image.
arXiv:2412.01106, 2024.
[12] Goodfellow IJ, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, Courville A, Bengio Y. Generative adversarial networks.
arXiv:1406.2661, 2014.
[13] Ho J, Jain A, Abbeel P. Denoising diffusion probabilistic models. arXiv:2006.11239, 2020.
[14] Isola P, Zhu JY, Zhou TH, Efros AA. Image-to-image translation with conditional adversarial networks. In: Proc. of the 2017 IEEE/CVF
Conf. on Computer Vision and Pattern Recognition (CVPR). Honolulu: IEEE, 2017. 5967–5976. [doi: 10.1109/CVPR.2017.632]
[15] Wang TC, Liu MY, Zhu JY, Tao A, Kautz J, Catanzaro B. High-resolution image synthesis and semantic manipulation with conditional
GANs. In: Proc. of the 2018 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). Salt Lake City: IEEE, 2018.
8798–8807. [doi: 10.1109/CVPR.2018.00917]
[16] Ma LQ, Jia X, Sun QR, Schiele B, Tuytelaars T, van Gool V. Pose guided person image generation. arXiv:1705.09368, 2018.
[17] Ronneberger O, Fischer P, Brox T. U-Net: Convolutional networks for biomedical image segmentation. In: Proc. of the 18th Int’l Conf.
on Medical Image Computing and Computer-assisted Intervention (MICCAI). Munich: Springer, 2015. 234–241. [doi: 10.1007/978-3-

