Page 441 - 《软件学报》2026年第4期
P. 441
1882 软件学报 2026 年第 37 卷第 4 期
[74] Mu JT, Gharbi M, Zhang R, Shechtman E, Vasconcelos N, Wang XL, Park T. Editable image elements for controllable synthesis. In:
Proc. of the 18th European Conf. on Computer Vision. Milan: Springer, 2024. 39–56. [doi: 10.1007/978-3-031-72627-9_3]
[75] Lee S, Gu G, Park S, Choi S, Choo J. High-resolution virtual try-on with misalignment and occlusion-handled conditions. In: Proc. of
the 17th European Conf. on Computer Vision. Milan: Springer, 2022. 204–219. [doi: 10.1007/978-3-031-19790-1_13]
[76] Morelli D, Baldrati A, Cartella G, Cornia M, Bertini M, Cucchiara R. LaDI-VTON: Latent diffusion textual-inversion enhanced virtual
try-on. In: Proc. of the 31st ACM Int’l Conf. on Multimedia. Ottawa: ACM, 2023. 8580–8589. [doi: 10.1145/3581783.3612137]
[77] Gou JH, Sun SY, Zhang JF, Si JL, Qian C, Zhang LQ. Taming the power of diffusion models for high-quality virtual try-on with
appearance flow. In: Proc. of the 31st ACM Int’l Conf. on Multimedia. Ottawa: ACM, 2023. 7599–7607. [doi: 10.1145/3581783.
3612255]
[78] Kim J, Gu G, Park M, Park S, Choo J. Stable VITON: Learning semantic correspondence with latent diffusion model for virtual try-on.
In: Proc. of the 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024. 8176–8185. [doi: 10.1109/
CVPR52733.2024.00781]
[79] Zhu LY, Yang DW, Zhu T, Reda F, Chan W, Saharia C, Norouzi M, Kemelmacher-Shlizerman I. TryOnDiffusion: A tale of two UNets.
In: Proc. of the 2023 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023. 4606–4615. [doi: 10.1109/
CVPR52729.2023.00447]
[80] Xu YH, Gu T, Chen WF, Chen A. OOTDiffusion: Outfitting fusion based latent diffusion for controllable virtual try-on. In: Proc. of the
39th AAAI Conf. on Artificial Intelligence. Philadelphia: AAAI Press, 2025. 8996–9004. [doi: 10.1609/aaai.v39i9.32973]
[81] Choi Y, Kwak S, Lee K, Choi H, Shin J. Improving diffusion models for authentic virtual try-on in the wild. In: Proc. of the 18th
European Conf. on Computer Vision. Milan: Springer, 2024. 206–235. [doi: 10.1007/978-3-031-73016-0_13]
[82] Zhang XP, Song D, Zhan PX, Chang TY, Zeng JH, Chen QG, Luo WH, Liu AA. BooW-VTON: Boosting in-the-wild virtual try-on via
mask-free pseudo data training. In: Proc. of the 2025 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Nashville: IEEE,
2025. 26399–26408. [doi: 10.1109/CVPR52734.2025.02458]
[83] Wang R, Guo HL, Liu JM, Li HX, Zhao HB, Tang X, Hu Y, Tang H, Li PP. StableGarment: Garment-centric generation via stable
diffusion. arXiv:2403.10783, 2024.
[84] Shen F, Jiang X, He X, Ye H, Wang C, Du XY, Li ZC, Tang JH. IMAGDressing-v1: Customizable virtual dressing. In: Proc. of the 39th
AAAI Conf. on Artificial Intelligence. Philadelphia: AAAI Press, 2025. 6795–6804. [doi: 10.1609/aaai.v39i7.32729]
[85] Chen WF, Gu T, Xu YH, Chen A. Magic Clothing: Controllable garment-driven image synthesis. In: Proc. of the 32nd ACM Int’l Conf.
on Multimedia. Melbourne: ACM, 2024. 6939–6948. [doi: 10.1145/3664647.3680691]
[86] Lin ET, Zhang XJ, Zhao FW, Luo YX, Dong X, Zeng L, Liang XD. DreamFit: Garment-centric human generation via a lightweight
anything-dressing encoder. In: Proc. of the 2025 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Nashville: IEEE, 2025.
23723–23733. [doi: 10.1109/CVPR52734.2025.02209]
[87] Li XH, Sun QC, Zhang PZ, Ye FL, Liao ZC, Feng WQ, Zhao ST, He Q. AnyDressing: Customizable multi-garment virtual dressing via
latent diffusion models. In: Proc. of the 39th AAAI Conf. on Artificial Intelligence. Philadelphia: AAAI Press, 2025. 5218–5226. [doi:
10.1609/aaai.v39i5.32554]
[88] Peng Y, Cui YX, Tang HM, Qi ZK, Dong RP, Bai J, Han CR, Ge Z, Zhang XY, Xia ST. DreamBench++: A human-aligned benchmark
for personalized image generation. In: Proc. of the 13th Int’l Conf. on Learning Representations. Singapore: OpenReview.net, 2025.
1–23.
[89] Wang WH, Lv QS, Yu WM, Hong WY, Qi J, Wang Y, Ji JH, Yang ZY, Zhao L, Song XX, Xu JZ, Chen KP, Xu B, Li JZ, Dong YX,
Ding M, Tang J. CogVLM: Visual expert for pretrained language models. In: Proc. of the 38th Int’l Conf. on Neural Information
Processing Systems. Vancouver: Curran Associates Inc., 2024. 3860.
[90] Liang J, Zeng H, Cui MM, Xie XS, Zhang L. PPR10K: A large-scale portrait photo retouching dataset with human-region mask and
group-level consistency. In: Proc. of the 2021 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Nashville: IEEE, 2021.
653–661. [doi: 10.1109/CVPR46437.2021.00071]
[91] Jafarian Y, Park HS. Self-supervised 3D representation learning of dressed humans from social media videos. IEEE Trans. on Pattern
Analysis and Machine Intelligence, 2022, 45(7): 8969–8983. [doi: 10.1109/TPAMI.2022.3231558]
[92] Liu ZW, Luo P, Wang XG, Tang XO. Deep learning face attributes in the wild. In: Proc. of the 2015 IEEE Int’l Conf. on Computer
Vision. Santiago: IEEE, 2015. 3730–3738. [doi: 10.1109/ICCV.2015.425]
[93] Choi S, Park S, Lee M, Choo J. VITON-HD: High-resolution virtual try-on via misalignment-aware normalization. In: Proc. of the 2021
IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Nashville: IEEE, 2021. 14126–14135. [doi: 10.1109/CVPR46437.2021.
01391]

