Page 121 - 《软件学报》2026年第5期
P. 121

2000                                                       软件学报  2026  年第  37  卷第  5  期


                      319-24574-4_28]
                 [18]   Esser P, Sutter E. A variational U-Net for conditional appearance and shape generation. In: Proc. of the 2018 IEEE/CVF Conf. on
                      Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018. 8857–8866. [doi: 10.1109/CVPR.2018.00923]
                 [19]   Bhunia AK, Khan S, Cholakkal H, Anwer RM, Laaksonen J, Shah M, Khan FS. Person image synthesis via denoising diffusion model.
                      In: Proc. of the 2023 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). Vancouver: IEEE, 2023. 5968–5976. [doi:
                      10.1109/cvpr52729.2023.00578]
                 [20]   Xu  LH,  Zhao  L,  Sun  XX,  Yan  KD,  Li  GY.  A  comprehensive  review  of  progress  in  deep-learning-based  occluded  human  pose
                      estimation. Journal of Image and Graphics, 2024, 29(12): 3529–3542 (in Chinese with English abstract). [doi: 10.11834/jig.230730]
                 [21]   Toshev A, Szegedy C. DeepPose: Human pose estimation via deep neural networks. In: Proc. of the 2014 IEEE Conf. on Computer
                      Vision and Pattern Recognition (CVPR). Columbus: IEEE, 2014. 1653–1660. [doi: 10.1109/cvpr.2014.214]
                 [22]   Cao Z, Hidalgo G, Simon T, Wei SE, Sheikh Y. OpenPose: Realtime multi-person 2D pose estimation using part affinity fields. IEEE
                      Trans. on Pattern Analysis and Machine Intelligence, 2021, 43(1): 172–186. [doi: 10.1109/TPAMI.2019.2929257]
                 [23]   Fang HS, Li JF, Tang HY, Xu C, Zhu HY, Xiu YL, Li YL, Lu CW. AlphaPose: Whole-body regional multi-person pose estimation and
                      tracking in real-time. IEEE Trans. on Pattern Analysis and Machine Intelligence, 2023, 45(6): 7157–7173. [doi: 10.1109/tpami.2022.
                      3222784]
                 [24]   Yang HH, Liu HX, Zhang YM, Wu XJ. Parallel multi-scale spatio-temporal graph convolutional network for 3D human pose estimation.
                      Ruan Jian Xue Bao/Journal of Software, 2025, 36(5): 2151–2166 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/
                      7200.htm [doi: 10.13328/j.cnki.jos.007200]
                 [25]   He JH, Sun JY, Liu Q. Multi-person 3D pose estimation using human-and-scene contexts. Ruan Jian Xue Bao/Journal of Software,
                      2024,  35(4):  2039–2054  (in  Chinese  with  English  abstract).  http://www.jos.org.cn/1000-9825/6837.htm  [doi:  10.13328/j.cnki.jos.
                      006837]
                 [26]   Güler RA, Neverova N, Kokkinos I. DensePose: Dense human pose estimation in the wild. In: Proc. of the 2018 IEEE/CVF Conf. on
                      Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018. 7297–7306. [doi: 10.1109/CVPR.2018.00762]
                 [27]   Loper M, Mahmood N, Romero J, Pons-Moll G, Black MJ. SMPL: A skinned multi-person linear model. Seminal Graphics Papers:
                      Pushing the Boundaries, 2023, 2: 851–866. [doi: 10.1145/3596711.3596800]
                 [28]   Romero J, Tzionas D, Black MJ. Embodied hands: Modeling and capturing hands and bodies together. ACM Trans. on Graphics, 2017,
                      36(6): 245. [doi: 10.1145/3130800.3130883]
                 [29]   Pavlakos G, Choutas V, Ghorbani N, Bolkart T, Osman AA, Tzionas D, Black MJ. Expressive body capture: 3D hands, face, and body
                      from a single image. In: Proc. of the 2019 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). Long Beach: IEEE,
                      2019. 10967–10977. [doi: 10.1109/cvpr.2019.01123]
                 [30]   Karras J, Holynski A, Wang TC, Kemelmacher-Shlizerman I. DreamPose: Fashion image-to-video synthesis via stable diffusion. In:
                      Proc. of the 2023 IEEE/CVF Int’l Conf. on Computer Vision (ICCV). Paris: IEEE, 2023. 22623–22633. [doi: 10.1109/iccv51070.2023.
                      02073]
                 [31]   Feng MY, Liu JL, Yu K, Yao Y, Hui Z, Guo XF, Lin XH, Xue HL, Shi C, Li XW, Li AJ, Kang XY, Lei BW, Cui MM, Ren PR, Xie
                      XS. DreaMoving: A human video generation framework based on diffusion models. arXiv:2312.05107, 2023.
                 [32]   Tu  SY,  Dai  Q,  Zhang  ZH,  Xie  SC,  Cheng  ZQ,  Luo  C,  Han  XT,  Wu  ZX,  Jiang  YG.  MotionFollower:  Editing  video  motion  via
                      lightweight score-guided diffusion. arXiv:2405.20325, 2024.
                 [33]   Radford  A,  Kim  JW,  Hallacy  C,  Ramesh  A,  Goh  G,  Agarwal  S,  Sastry  G,  Askell  A,  Mishkin  P,  Clark  J,  Krueger  G,  Sutskever  I.
                      Learning transferable visual models from natural language supervision. arXiv:2103.00020, 2021.
                 [34]   Chang D, Shi YC, Gao QK, Fu J, Xu HY, Song GX, Yan Q, Zhu YZ, Yang X, Soleymani M. MagicPose: Realistic human poses and
                      facial expressions retargeting with identity-aware diffusion. arXiv:2311.12052, 2024.
                 [35]   Xu ZC, Zhang JF, Liew JH, Yan HS, Liu JW, Zhang CX, Feng JS, Shou MZ. MagicAnimate: Temporally consistent human image
                      animation using diffusion model. In: Proc. of the 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). Seattle:
                      IEEE, 2024. 1481–1490. [doi: 10.1109/CVPR52733.2024.00147]
                 [36]   Li H. Animate Anyone: Consistent and controllable image-to-video synthesis for character animation. In: Proc. of the 2024 IEEE/CVF
                      Conf. on Computer Vision and Pattern Recognition (CVPR). Seattle: IEEE, 2024. 8153–8163. [doi: 10.1109/cvpr52733.2024.00779]
                 [37]   Liu JL, Yu K, Feng MY, Guo XF, Cui MM. Disentangling foreground and background motion for enhanced realism in human video
                      generation. arXiv:2405.16393, 2024.
                 [38]   Ma LQ, Sun QR, Georgoulis S, van Gool V, Schiele B, Fritz M. Disentangled person image generation. In: Proc. of the 2018 IEEE/CVF
                      Conf. on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018. 99–108. [doi: 10.1109/cvpr.2018.00018]
   116   117   118   119   120   121   122   123   124   125   126