Page 126 - 《软件学报》2026年第5期
P. 126

李玘芮 等: 姿态控制人物图像视频生成技术综述                                                         2005


                 [128]   Liu YF, Cun XD, Liu XB, Wang XT, Zhang Y, Chen HX, Liu Y, Zeng TY, Chan R, Shan Y. EvalCrafter: Benchmarking and evaluating
                      large video generation models. In: Proc. of the 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). Seattle:
                      IEEE, 2024. 22139–22149. [doi: 10.1109/CVPR52733.2024.02090]
                 [129]   Yoon Y, Cha B, Lee JH, Jang M, Lee J, Kim J, Lee G. Speech gesture generation from the trimodal context of text, audio, and speaker
                      identity. ACM Trans. on Graphics, 2020, 39(6): 222. [doi: 10.1145/3414685.3417838]
                 [130]   Salimans T, Goodfellow I, Zaremba W, Cheung V, Radford A, Chen X. Improved techniques for training GANs. In: Proc. of the 30th
                      Int’l Conf. on Neural Information Processing Systems. Barcelona: Curran Associates Inc., 2016. 2234–2242.
                 [131]   Liu X, Wu QY, Zhou H, Xu YH, Qian R, Lin XY, Zhou XW, Wu W, Dai B, Zhou BL. Learning hierarchical cross-modal association
                      for co-speech gesture generation. In: Proc. of the 2022 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). New
                      Orleans: IEEE, 2022. 10452–10462. [doi: 10.1109/CVPR52688.2022.01021]
                 [132]   Wu HN, Zhang EL, Liao L, Chen CF, Hou JW, Wang AN, Sun WX, Yan Q, Lin WS. Exploring video quality assessment on user
                      generated contents from aesthetic and technical perspectives. In: Proc. of the 2023 IEEE/CVF Int’l Conf. on Computer Vision (ICCV).
                      Paris: IEEE, 2023. 20087–20097. [doi: 10.1109/ICCV51070.2023.01843]
                 [133]   Yang  Y,  Ramanan  D.  Articulated  human  detection  with  flexible  mixtures  of  parts.  IEEE  Trans.  on  Pattern  Analysis  and  Machine
                      Intelligence, 2013, 35(12): 2878–2890. [doi: 10.1109/TPAMI.2012.261]
                 [134]   Esser P, Chiu J, Atighehchian P, Granskog J, Germanidis A. Structure and content-guided video synthesis with diffusion models. In:
                      Proc. of the 2023 IEEE/CVF Int’l Conf. on Computer Vision (ICCV). Paris: IEEE, 2023. 7312–7322. [doi: 10.1109/ICCV51070.2023.
                      00675]
                 [135]   Chen HH, Tao YF, Zhang JY. A survey on 3D clothed human body reconstruction: From traditional methods to high-fidelity models.
                      Journal of Image and Graphics, 2024, 29(9): 2566–2595 (in Chinese with English abstract). [doi: 10.11834/jig.230646]
                 [136]   Xu  YH,  Gu  T,  Chen  WF,  Chen  CC.  OOTDiffusion:  Outfitting  fusion  based  latent  diffusion  for  controllable  virtual  try-on.
                      arXiv:2403.01779, 2024.
                 [137]   Cheong SY, Mustafa A, Gilbert A. UPGPT: Universal diffusion model for person image generation, editing and pose transfer. In: Proc.
                      of the 2023 IEEE/CVF Int’l Conf. on Computer Vision Workshops (ICCVW). Paris: IEEE, 2023. 4175–4184. [doi: 10.1109/ICCVW
                      60793.2023.00451]

                 附中文参考文献
                 [20]   徐琳皓, 赵林, 孙辛欣, 颜克冬, 李广宇. 基于深度学习的遮挡人体姿态估计进展综述. 中国图象图形学报, 2024, 29(12):
                      3529–3542. [doi: 10.11834/jig.230730]
                 [24]   杨红红, 刘泓希, 张玉梅, 吴晓军. 基于平行多尺度时空图卷积网络的三维人体姿态估计算法. 软件学报, 2025, 36(5): 2151–2166.
                      http://www.jos.org.cn/1000-9825/7200.htm [doi: 10.13328/j.cnki.jos.007200]
                 [25]   何建航, 孙郡瑤, 刘琼. 基于人体和场景上下文的多人     3D  姿态估计. 软件学报, 2024, 35(4): 2039–2054. http://www.jos.org.cn/1000-
                      9825/6837.htm [doi: 10.13328/j.cnki.jos.006837]
                 [135]   陈鸿鹄, 陶云帆, 张举勇. 三维穿衣人体重建综述——从传统方法到高保真模型. 中国图象图形学报, 2024, 29(9): 2566–2595. [doi:
                      10.11834/jig.230646]

                 作者简介
                 李玘芮, 硕士生, CCF  学生会员, 主要研究领域为人工智能生成, 推理加速.
                 励雪巍, 博士, 特聘副研究员, CCF  专业会员, 主要研究领域为计算机视觉, 人工智能, 如知识蒸馏, 场景图生成, 扩散模型.
                 赵奇, 硕士生, 主要研究领域为人工智能生成, 推理加速.
                 李杰, 硕士生, 主要研究领域为人工智能生成, 推理加速.
                 李玺, 博士, 教授, 博士生导师, CCF  杰出会员, 主要研究领域为人工智能, 计算机视觉, 机器学习, 模式识别, 数据挖掘.
   121   122   123   124   125   126   127   128   129   130   131