Page 125 - 《软件学报》2026年第5期
P. 125
2004 软件学报 2026 年第 37 卷第 5 期
[105] Fu XM, Wang X, Liu J, Wang SH, Dai J, Han JZ. CoP: Chain-of-pose for image animation in large pose changes. In: Proc. of the 31st
ACM Int’l Conf. on Multimedia. Ottawa: ACM, 2023. 6969–6977. [doi: 10.1145/3581783.3611772]
[106] Zheng L, Shen LY, Tian L, Wang SJ, Wang JD, Tian Q. Scalable person re-identification: A benchmark. In: Proc. of the 2015 IEEE Int’l
Conf. on Computer Vision (ICCV). Santiago: IEEE, 2015. 1116–1124. [doi: 10.1109/iccv.2015.133]
[107] Liu ZW, Luo P, Qiu S, Wang XG, Tang XO. DeepFashion: Powering robust clothes recognition and retrieval with rich annotations. In:
Proc. of the 2016 IEEE Conf. on Computer Vision and Pattern Recognition (CVPR). Las Vegas: IEEE, 2016. 1096–1104. [doi: 10.1109/
CVPR.2016.124]
[108] Liu KH, Chen TY, Chen CS. MVC: A dataset for view-invariant clothing retrieval and attribute prediction. In: Proc. of the 2016 ACM
on Int’ l Conf. on Multimedia Retrieval. New York: ACM, 2016. 313–316. [doi: 10.1145/2911996.2912058]
[109] Yoo D, Kim N, Park S, Paek AS, Kweon IS. Pixel-level domain transfer. In: Proc. of the 14th European Conf. on Computer Vision.
Amsterdam: Springer, 2016. 517–532. [doi: 10.1007/978-3-319-46484-8_31]
[110] Dong HY, Liang XD, Shen XH, Wang BC, Lai HJ, Zhu J, Hu ZT, Yin J. Towards multi-pose guided virtual try-on network. In: Proc. of
the 2019 IEEE/CVF Int’l Conf. on Computer Vision (ICCV). Seoul: IEEE, 2019. 9025–9034. [doi: 10.1109/ICCV.2019.00912]
[111] Hsieh CW, Chen CY, Chou CL, Shuai HH, Liu JY, Cheng WH. FashionOn: Semantic-guided image-based virtual try-on with detailed
human and clothing information. In: Proc. of the 27th ACM Int’l Conf. on Multimedia. Nice: ACM, 2019. 275–283. [doi: 10.1145/
3343031.3351075]
[112] Zheng N, Song XM, Chen ZZ, Hu LM, Cao D, Nie LQ. Virtually trying on new clothing with arbitrary poses. In: Proc. of the 27th ACM
Int’l Conf. on Multimedia. Nice: ACM, 2019. 266–274. [doi: 10.1145/3343031.3350946]
[113] Jiang YM, Yang S, Qiu HN, Wu W, Loy CC, Liu ZW. Text2Human: Text-driven controllable human image generation. ACM Trans. on
Graphics, 2022, 41(4): 162. [doi: 10.1145/3528223.3530104]
[114] Fu JL, Li SK, Jiang YM, Lin KY, Qian C, Loy CC, Wu W, Liu ZW. StyleGAN-human: A data-centric odyssey of human generation. In:
Proc. of the 17th European Conf. on Computer Vision. Tel Aviv: Springer, 2022. 1–19. [doi: 10.1007/978-3-031-19787-1_1]
[115] Soomro K, Zamir AR, Shah M. UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv:1212.0402, 2012.
[116] Zhang WY, Zhu ML, Derpanis KG. From actemes to action: A strongly-supervised representation for detailed action understanding. In:
Proc. of the 2013 IEEE Int’l Conf. on Computer Vision. Sydney: IEEE, 2013. 2248–2255. [doi: 10.1109/ICCV.2013.280]
[117] Ionescu C, Papava D, Olaru V, Sminchisescu C. Human3.6M: Large scale datasets and predictive methods for 3D human sensing in
natural environments. IEEE Trans. on Pattern Analysis and Machine Intelligence, 2014, 36(7): 1325–1339. [doi: 10.1109/tpami.2013.
248]
[118] Wang BY, Wang XF, Ni CJ, Zhao GS, Yang ZQ, Zhu Z, Zhang MY, Zhou YK, Chen XZ, Huang G, Liu LH, Wang XG.
HumanDreamer: Generating controllable human-motion videos via decoupled generation. arXiv:2503.24026, 2025.
[119] Dong HY, Liang XD, Shen XH, Wu BW, Chen BC, Yin J. FW-GAN: Flow-navigated warping GAN for video virtual try-on. In: Proc.
of the 2019 IEEE/CVF Int’l Conf. on Computer Vision (ICCV). Seoul: IEEE, 2019. 1161–1170. [doi: 10.1109/iccv.2019.00125]
[120] Jafarian Y, Park HS. Self-supervised 3D representation learning of dressed humans from social media videos. IEEE Trans. on Pattern
Analysis and Machine Intelligence, 2023, 45(7): 8969–8983. [doi: 10.1109/TPAMI.2022.3231558]
[121] Bishop CM. Pattern Recognition and Machine Learning. New York: Springer, 2006. 47.
[122] Horé A, Ziou D. Image quality metrics: PSNR vs. SSIM. In: Proc. of the 20th Int’l Conf. on Pattern Recognition. Istanbul: IEEE, 2010.
2366–2369. [doi: 10.1109/ICPR.2010.579]
[123] Wang Z, Bovik AC, Sheikh HR, Simoncelli EP. Image quality assessment: From error visibility to structural similarity. IEEE Trans. on
Image Processing, 2004, 13(4): 600–612. [doi: 10.1109/TIP.2003.819861]
[124] Zhang R, Isola P, Efros AA, Shechtman E, Wang O. The unreasonable effectiveness of deep features as a perceptual metric. In: Proc. of
the 2018 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018. 586–595. [doi: 10.1109/CVPR.
2018.00068]
[125] Rabin J, Peyré G, Delon J, Bernot M. Wasserstein barycenter and its application to texture mixing. In: Proc. of the 3rd Int’l Conf. on
Scale Space and Variational Methods in Computer Vision. Ein-Gedi: Springer, 2011. 435–446. [doi: 10.1007/978-3-642-24785-9_37]
[126] Heusel M, Ramsauer H, Unterthiner T, Nessler B, Hochreiter S. GANs trained by a two time-scale update rule converge to a local nash
equilibrium. In: Proc. of the 31st Int’l Conf. on Neural Information Processing Systems. Long Beach: Curran Associates Inc., 2017.
6629–6640.
[127] Tulyakov S, Liu MY, Yang XD, Kautz J. MoCoGAN: Decomposing motion and content for video generation. In: Proc. of the 2018
IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018. 1526–1535. [doi: 10.1109/CVPR.2018.
00165]

