Page 124 - 《软件学报》2026年第5期
P. 124
李玘芮 等: 姿态控制人物图像视频生成技术综述 2003
[82] Siarohin A, Lathuilière S, Tulyakov S, Ricci E, Sebe N. First order motion model for image animation. arXiv:2003.00196, 2020.
[83] Wang BC, Zheng HB, Liang XD, Chen YM, Lin L, Yang M. Toward characteristic-preserving image-based virtual try-on network. In:
Proc. of the 15th European Conf. on Computer Vision. Munich: Springer, 2018. 607–623. [doi: 10.1007/978-3-030-01261-8_36]
[84] Wang QL, Jiang ZK, Xu CM, Zhang JN, Wang YB, Zhang XY, Cao Y, Cao WJ, Wang CJ, Fu YW. VividPose: Advancing stable video
diffusion for realistic human image animation. arXiv:2405.18156, 2024.
[85] Wang TC, Liu MY, Tao A, Liu GL, Kautz J, Catanzaro B. Few-shot video-to-video synthesis. In: Proc. of the 33rd Int’l Conf. on Neural
Information Processing Systems. Vancouver: Curran Associates Inc., 2019. 5013–5024.
[86] Yang ZD, Zeng AL, Yuan C, Li Y. Effective whole-body pose estimation with two-stages distillation. arXiv:2307.15880, 2023.
[87] Wang TC, Liu MY, Zhu JY, Liu GL, Tao A, Kautz J, Catanzaro B. Video-to-video synthesis. In: Proc. of the 32nd Int’l Conf. on Neural
Information Processing Systems. Montréal: Curran Associates Inc., 2018. 1152–1164.
[88] Shao RZ, Pang YX, Zheng ZR, Sun JX, Liu YB. Human4DiT: 360-degree human video generation with 4D diffusion Transformer.
arXiv:2405.17405, 2024.
[89] Zhu SH, Chen JL, Dai ZZ, Dong ZL, Xu YH, Cao X, Yao Y, Zhu H, Zhu SY. CHAMP: Controllable and consistent human image
animation with 3D parametric guidance. In: Proc. of the 18th European Conf. on Computer Vision. Milan: Springer, 2025. 145–162.
[doi: 10.1007/978-3-031-73001-6_9]
[90] Zhang LM, Rao AY, Agrawala M. Adding conditional control to text-to-image diffusion models. In: Proc. of the 2023 IEEE/CVF Int’l
Conf. on Computer Vision (ICCV). Paris: IEEE, 2023. 3813–3824. [doi: 10.1109/iccv51070.2023.00355]
[91] Yan YC, Xu JW, Ni BB, Zhang WD, Yang XK. Skeleton-aided articulated motion generation. In: Proc. of the 25th ACM Int’l Conf. on
Multimedia. Mountain View: ACM, 2017. 199–207. [doi: 10.1145/3123266.3123277]
[92] Guo YW, Yang CY, Rao AY, Liang ZY, Wang YH, Qiao Y, Agrawala M, Lin DH, Dai B. AnimateDiff: Animate your personalized text-
to-image diffusion models without specific tuning. arXiv:2307.04725, 2024.
[93] Gan QJ, Ren Y, Zhang C, Ye ZH, Xie P, Yin X, Yuan ZH, Peng BY, Zhu JK. HumanDiT: Pose-guided diffusion Transformer for long-
form human motion video generation. arXiv:2502.04847, 2025.
[94] Pumarola A, Agudo A, Sanfeliu A, Moreno-Noguer F. Unsupervised person image synthesis in arbitrary poses. In: Proc. of the 2018
IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018. 8620–8628. [doi: 10.1109/cvpr.2018.
00899]
[95] Tang H, Xu D, Liu GW, Wang W, Sebe N, Yan Y. Cycle in cycle generative adversarial networks for keypoint-guided image
generation. In: Proc. of the 27th ACM Int’l Conf. on Multimedia. Nice: ACM, 2019. 2052–2060. [doi: 10.1145/3343031.3350980]
[96] Song SJ, Zhang W, Liu JY, Mei T. Unsupervised person image generation with semantic parsing transformation. In: Proc. of the 2019
IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). Long Beach: IEEE, 2019. 2352–2361. [doi: 10.1109/CVPR.
2019.00246]
[97] Ma Y, He YQ, Cun XD, Wang XT, Chen SR, Li X, Chen QF. Follow your pose: Pose-guided text-to-video generation using pose-free
videos. In: Proc. of the 38th AAAI Conf. on Artificial Intelligence. Vancouver: AAAI, 2024. 4117–4125. [doi: 10.1609/aaai.v38i5.
28206]
[98] Zhang PZ, Yang LX, Lai JH, Xie XH. Exploring dual-task correlation for pose guided person image generation. In: Proc. of the 2022
IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). New Orleans: IEEE, 2022. 7703–7712. [doi: 10.1109/CVPR
52688.2022.00756]
[99] Shen F, Ye H, Zhang J, Wang C, Han X, Yang W. Advancing pose-guided image synthesis with progressive conditional diffusion
models. arXiv:2310.06313, 2024.
[100] Roy P, Bhattacharya S, Ghosh S, Pal U. Multi-scale attention guided pose transfer. Pattern Recognition, 2023, 137: 109315. [doi: 10.
1016/j.patcog.2023.109315]
[101] Yu WY, Po LM, Zhao YZ, Xiong JJ, Lau KW. Spatial content alignment for pose transfer. In: Proc. of the 2021 IEEE Int’l Conf. on
Multimedia and Expo (ICME). Shenzhen: IEEE, 2021. 1–6. [doi: 10.1109/ICME51207.2021.9428146]
[102] Lu YZ, Zhang ML, Ma AJ, Xie XH, Lai JH. Coarse-to-fine latent diffusion for pose-guided person image synthesis. In: Proc. of the
2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). Seattle: IEEE, 2024. 6420–6429. [doi: 10.1109/CVPR
52733.2024.00614]
[103] Barnes C, Shechtman E, Finkelstein A, Goldman DB. PatchMatch: A randomized correspondence algorithm for structural image editing.
ACM Trans. on Graphics, 2009, 28(3): 24. [doi: 10.1145/1531326.1531330]
[104] Wang X, Zhang SW, Tang LX, Zhang YY, Gao CX, Wang YH, Sang N. UniAnimate-DiT: Human image animation with large-scale
video diffusion Transformer. arXiv:2504.11289, 2025.

