Page 121 - 《软件学报》2026年第5期
P. 121
2000 软件学报 2026 年第 37 卷第 5 期
319-24574-4_28]
[18] Esser P, Sutter E. A variational U-Net for conditional appearance and shape generation. In: Proc. of the 2018 IEEE/CVF Conf. on
Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018. 8857–8866. [doi: 10.1109/CVPR.2018.00923]
[19] Bhunia AK, Khan S, Cholakkal H, Anwer RM, Laaksonen J, Shah M, Khan FS. Person image synthesis via denoising diffusion model.
In: Proc. of the 2023 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). Vancouver: IEEE, 2023. 5968–5976. [doi:
10.1109/cvpr52729.2023.00578]
[20] Xu LH, Zhao L, Sun XX, Yan KD, Li GY. A comprehensive review of progress in deep-learning-based occluded human pose
estimation. Journal of Image and Graphics, 2024, 29(12): 3529–3542 (in Chinese with English abstract). [doi: 10.11834/jig.230730]
[21] Toshev A, Szegedy C. DeepPose: Human pose estimation via deep neural networks. In: Proc. of the 2014 IEEE Conf. on Computer
Vision and Pattern Recognition (CVPR). Columbus: IEEE, 2014. 1653–1660. [doi: 10.1109/cvpr.2014.214]
[22] Cao Z, Hidalgo G, Simon T, Wei SE, Sheikh Y. OpenPose: Realtime multi-person 2D pose estimation using part affinity fields. IEEE
Trans. on Pattern Analysis and Machine Intelligence, 2021, 43(1): 172–186. [doi: 10.1109/TPAMI.2019.2929257]
[23] Fang HS, Li JF, Tang HY, Xu C, Zhu HY, Xiu YL, Li YL, Lu CW. AlphaPose: Whole-body regional multi-person pose estimation and
tracking in real-time. IEEE Trans. on Pattern Analysis and Machine Intelligence, 2023, 45(6): 7157–7173. [doi: 10.1109/tpami.2022.
3222784]
[24] Yang HH, Liu HX, Zhang YM, Wu XJ. Parallel multi-scale spatio-temporal graph convolutional network for 3D human pose estimation.
Ruan Jian Xue Bao/Journal of Software, 2025, 36(5): 2151–2166 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/
7200.htm [doi: 10.13328/j.cnki.jos.007200]
[25] He JH, Sun JY, Liu Q. Multi-person 3D pose estimation using human-and-scene contexts. Ruan Jian Xue Bao/Journal of Software,
2024, 35(4): 2039–2054 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/6837.htm [doi: 10.13328/j.cnki.jos.
006837]
[26] Güler RA, Neverova N, Kokkinos I. DensePose: Dense human pose estimation in the wild. In: Proc. of the 2018 IEEE/CVF Conf. on
Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018. 7297–7306. [doi: 10.1109/CVPR.2018.00762]
[27] Loper M, Mahmood N, Romero J, Pons-Moll G, Black MJ. SMPL: A skinned multi-person linear model. Seminal Graphics Papers:
Pushing the Boundaries, 2023, 2: 851–866. [doi: 10.1145/3596711.3596800]
[28] Romero J, Tzionas D, Black MJ. Embodied hands: Modeling and capturing hands and bodies together. ACM Trans. on Graphics, 2017,
36(6): 245. [doi: 10.1145/3130800.3130883]
[29] Pavlakos G, Choutas V, Ghorbani N, Bolkart T, Osman AA, Tzionas D, Black MJ. Expressive body capture: 3D hands, face, and body
from a single image. In: Proc. of the 2019 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). Long Beach: IEEE,
2019. 10967–10977. [doi: 10.1109/cvpr.2019.01123]
[30] Karras J, Holynski A, Wang TC, Kemelmacher-Shlizerman I. DreamPose: Fashion image-to-video synthesis via stable diffusion. In:
Proc. of the 2023 IEEE/CVF Int’l Conf. on Computer Vision (ICCV). Paris: IEEE, 2023. 22623–22633. [doi: 10.1109/iccv51070.2023.
02073]
[31] Feng MY, Liu JL, Yu K, Yao Y, Hui Z, Guo XF, Lin XH, Xue HL, Shi C, Li XW, Li AJ, Kang XY, Lei BW, Cui MM, Ren PR, Xie
XS. DreaMoving: A human video generation framework based on diffusion models. arXiv:2312.05107, 2023.
[32] Tu SY, Dai Q, Zhang ZH, Xie SC, Cheng ZQ, Luo C, Han XT, Wu ZX, Jiang YG. MotionFollower: Editing video motion via
lightweight score-guided diffusion. arXiv:2405.20325, 2024.
[33] Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, Krueger G, Sutskever I.
Learning transferable visual models from natural language supervision. arXiv:2103.00020, 2021.
[34] Chang D, Shi YC, Gao QK, Fu J, Xu HY, Song GX, Yan Q, Zhu YZ, Yang X, Soleymani M. MagicPose: Realistic human poses and
facial expressions retargeting with identity-aware diffusion. arXiv:2311.12052, 2024.
[35] Xu ZC, Zhang JF, Liew JH, Yan HS, Liu JW, Zhang CX, Feng JS, Shou MZ. MagicAnimate: Temporally consistent human image
animation using diffusion model. In: Proc. of the 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). Seattle:
IEEE, 2024. 1481–1490. [doi: 10.1109/CVPR52733.2024.00147]
[36] Li H. Animate Anyone: Consistent and controllable image-to-video synthesis for character animation. In: Proc. of the 2024 IEEE/CVF
Conf. on Computer Vision and Pattern Recognition (CVPR). Seattle: IEEE, 2024. 8153–8163. [doi: 10.1109/cvpr52733.2024.00779]
[37] Liu JL, Yu K, Feng MY, Guo XF, Cui MM. Disentangling foreground and background motion for enhanced realism in human video
generation. arXiv:2405.16393, 2024.
[38] Ma LQ, Sun QR, Georgoulis S, van Gool V, Schiele B, Fritz M. Disentangled person image generation. In: Proc. of the 2018 IEEE/CVF
Conf. on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018. 99–108. [doi: 10.1109/cvpr.2018.00018]

