Page 122 - 《软件学报》2026年第5期
P. 122
李玘芮 等: 姿态控制人物图像视频生成技术综述 2001
[39] Liu W, Piao ZX, Min J, Luo WH, Ma L, Gao SH. Liquid warping GAN: A unified framework for human motion imitation, appearance
transfer and novel view synthesis. In: Proc. of the 2019 IEEE/CVF Int’l Conf. on Computer Vision (ICCV). Seoul: IEEE, 2019.
5903–5912. [doi: 10.1109/ICCV.2019.00600]
[40] Li H, Wang GY, Shen Z, Gao X, Meng DC, Zhuo L, Zhang P, Zhang B, Bo LF. Animate Anyone 2: High-fidelity character image
animation with environment affordance. arXiv:2502.06145, 2025.
[41] Yang LB, Zhao ZH, Wang SQ, Wang SS, Ma SW, Gao W. Disentangled human action video generation via decoupled learning. In:
Proc. of the 2019 IEEE Int’l Conf. on Multimedia & Expo Workshops. Shanghai: IEEE, 2019. 495–500. [doi: 10.1109/icmew.2019.
00091]
[42] Li NN, Shih KJ, Plummer BA. Collecting the puzzle pieces: Disentangled self-driven human pose transfer by permuting textures. In:
Proc. of the 2023 IEEE/CVF Int’l Conf. on Computer Vision (ICCV). Paris: IEEE, 2023. 7092–7103. [doi: 10.1109/ICCV51070.2023.
00655]
[43] Balakrishnan G, Zhao A, Dalca AV, Durand F, Guttag J. Synthesizing images of humans in unseen poses. In: Proc. of the 2018
IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). Salt Lake City: IEEE, 2018. 8340–8348. [doi: 10.1109/cvpr.
2018.00870]
[44] Men YF, Mao YM, Jiang YN, Ma WY, Lian ZH. Controllable person image synthesis with attribute-decomposed GAN. In: Proc. of the
2020 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). Seattle: IEEE, 2020. 5083–5092. [doi: 10.1109/
CVPR42600.2020.00513]
[45] Pu G, Men Y, Mao YM, Jiang YN, Ma WY, Lian ZH. Controllable image synthesis with attribute-decomposed GAN. IEEE Trans. on
Pattern Analysis and Machine Intelligence, 2023, 45(2): 1514–1532. [doi: 10.1109/tpami.2022.3161985]
[46] Zhang JS, Li K, Lai YK, Yang JY. PISE: Person image synthesis and editing with decoupled GAN. In: Proc. of the 2021 IEEE/CVF
Conf. on Computer Vision and Pattern Recognition (CVPR). Nashville: IEEE, 2021. 7978–7986. [doi: 10.1109/cvpr46437.2021.00789]
[47] Ren YR, Fan XQ, Li G, Liu S, Li TH. Neural texture extraction and distribution for controllable person image synthesis. In: Proc. of the
2022 IEEE/CVF Conf. Computer Vision and Pattern Recognition (CVPR). New Orleans: IEEE, 2022. 13525–13534. [doi: 10.1109/
CVPR52688.2022.01317]
[48] Yang LB, Wang P, Liu C, Gao ZN, Ren PR, Zhang XF, Wang SS, Ma SW, Hua XS, Gao W. Towards fine-grained human pose transfer
with detail replenishing network. IEEE Trans. on Image Processing, 2021, 30: 2422–2435. [doi: 10.1109/tip.2021.3052364]
[49] Chan C, Ginosar S, Zhou TH, Efros A. Everybody dance now. In: Proc. of the 2019 IEEE/CVF Int’l Conf. on Computer Vision (ICCV).
Seoul: IEEE, 2019. 5932–5941. [doi: 10.1109/iccv.2019.00603]
[50] Huang SY, Xiong HY, Cheng ZQ, Wang QZ, Zhou XR, Wen BH, Huan J, Dou DJ. Generating person images with appearance-aware
pose stylizer. In: Proc. of the 29th Int’l Joint Conf. on Artificial Intelligence (IJCAI). Yokohama: Morgan Kaufmann, 2021. 623–629.
[51] Grigorev A, Sevastopolsky A, Vakhitov A, Lempitsky V. Coordinate-based texture inpainting for pose-guided human image generation.
In: Proc. of the 2019 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). Long Beach: IEEE, 2019. 12127–12136.
[doi: 10.1109/cvpr.2019.01241]
[52] Li YN, Huang C, Loy CC. Dense intrinsic appearance flow for human pose transfer. In: Proc. of the 2019 IEEE/CVF Conf. on Computer
Vision and Pattern Recognition (CVPR). Long Beach: IEEE, 2019. 3688–3697. [doi: 10.1109/cvpr.2019.00381]
[53] Zanfir M, Popa AI, Zanfir A, Sminchisescu C. Human appearance transfer. In: Proc. of the 2018 IEEE/CVF Conf. on Computer Vision
and Pattern Recognition. Salt Lake City: IEEE, 2018. 5391–5399. [doi: 10.1109/CVPR.2018.00565]
[54] Neverova N, Güler RA, Kokkinos I. Dense pose transfer. In: Proc. of the 15th European Conf. on Computer Vision. Munich: Springer,
2018. 128–143. [doi: 10.1007/978-3-030-01219-9_8]
[55] Zhang P, Zhang B, Chen D, Yuan L, Wen F. Cross-domain correspondence learning for exemplar-based image translation. In: Proc. of
the 2020 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). Seattle: IEEE, 2020. 5142–5152. [doi: 10.1109/
CVPR42600.2020.00519]
[56] Zhou XR, Zhang B, Zhang T, Zhang P, Bao JM, Chen D, Zhang ZF, Wen F. CoCosNet v2: Full-resolution correspondence learning for
image translation. In: Proc. of the 2021 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). Nashville: IEEE, 2021.
11460–11470. [doi: 10.1109/CVPR46437.2021.01130]
[57] Wu ZH, Lin GS, Tao QY, Cai JF. M2E-try on Net: Fashion from model to everyone. In: Proc. of the 27th ACM Int’l Conf. on
Multimedia. Nice: ACM, 2019. 293–301. [doi: 10.1145/3343031.3351083]
[58] Jaderberg M, Simonyan K, Zisserman A, Kavukcuoglu K. Spatial Transformer networks. arXiv:1506.02025, 2016.
[59] Siarohin A, Sangineto E, Lathuilière S, Sebe N. Deformable GANs for pose-based human image generation. In: Proc. of the 2018
IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018. 3408–3416. [doi: 10.1109/cvpr.2018.

