Page 439 - 《软件学报》2026年第4期
P. 439
1880 软件学报 2026 年第 37 卷第 4 期
[32] Wei ZC, Su QK, Qin L, Wang WZ. MM-Diff: High-fidelity image personalization via multi-modal condition integration. arXiv:2403.
15059, 2024.
[33] He JJ, Geng YF, Bo LF. UniPortrait: A unified framework for identity-preserving single-and multi-human image personalization. arXiv:
2408.05939, 2024.
[34] Xiao GX, Yin TW, Freeman WT, Durand F, Han S. FastComposer: Tuning-free multi-subject image generation with localized attention.
Int’l Journal of Computer Vision, 2025, 133(3): 1175–1194. [doi: 10.1007/s11263-024-02227-z]
[35] Zhou ZG, Li J, Li HX, Chen N, Tang X. StoryMaker: Towards holistic consistent characters in text-to-image generation. arXiv:
2409.12576, 2024.
[36] Zhao J, Zheng HL, Wang CY, Lan L, Huang WR, Tang YH. MagicNaming: Consistent identity generation by finding a “name space” in
T2I diffusion models. In: Proc. of the 39th AAAI Conf. on Artificial Intelligence. Philadelphia: AAAI Press, 2024. 10439–10447. [doi:
10.1609/aaai.v39i10.33133]
[37] Nam J, Son S, Xu Z, Shi J, Liu DF, Liu F, Kim K, Zhou Y. Visual Persona: Foundation model for full-body human customization. In:
Proc. of the 2025 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Nashville: IEEE, 2025. 18630–18641. [doi: 10.1109/
CVPR52734.2025.01736]
[38] Li H. Animate Anyone: Consistent and controllable image-to-video synthesis for character animation. In: Proc. of the 2024 IEEE/CVF
Conf. on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024. 8153–8163. [doi: 10.1109/CVPR52733.2024.00779]
[39] Zhang Y, Gu JX, Wang LW, Wang H, Cheng JQ, Zhu YF, Zou FY. MimicMotion: High-quality human motion video generation with
confidence-aware pose guidance. arXiv:2406.19680, 2024.
[40] Wang X, Zhang SW, Gao CX, Wang JY, Zhou XQ, Zhang YY, Yan LX, Sang N. UniAnimate: Taming unified video diffusion models
for consistent human image animation. arXiv:2406.01188, 2024.
[41] Wang T, Li LJ, Lin K, Zhai YH, Lin CC, Yang ZY, Zhang HW, Liu ZC, Wang LJ. DisCo: Disentangled control for realistic human
dance generation. In: Proc. of the 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024. 9326–9336.
[doi: 10.1109/CVPR52733.2024.00891]
[42] Xu ZC, Zhang JF, Liew JH, Yan HS, Liu JW, Zhang CX, Feng JS, Shou MZ. MagicAnimate: Temporally consistent human image
animation using diffusion model. In: Proc. of the 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Seattle: IEEE,
2024. 1481–1490. [doi: 10.1109/CVPR52733.2024.00147]
[43] Zhu SH, Chen JL, Dai ZZ, Dong ZL, Xu YH, Cao X, Yao Y, Zhu H, Zhu SY. Champ: Controllable and consistent human image
animation with 3D parametric guidance. In: Proc. of the 18th European Conf. on Computer Vision. Milan: Springer, 2024. 145–162.
[doi: 10.1007/978-3-031-73001-6_9]
[44] Parihar R, Gupta H, Vs S, Babu RV. Text2Place: Affordance-aware text guided human placement. In: Proc. of the 18th European Conf.
on Computer Vision. Milan: Springer, 2024. 57–77. [doi: 10.1007/978-3-031-72646-0_4]
[45] Saini N, Bodla N, Shrivastava A, Ravichandran A, Zhang X, Shrivastava A, Singh B. InVi: Object insertion in videos using off-the-shelf
diffusion models. arXiv:2407.10958, 2024.
[46] Qiu D, Chen ZY, Wang R, Fan MY, Yu CQ, Huang JS, Wen X. MovieCharacter: A tuning-free framework for controllable character
video synthesis. arXiv:2410.20974, 2024.
[47] Xu ZY, Huang ZY, Cao J, Zhang Y, Cun XD, Shuai Q, Wang YC, Bao LC, Li JT, Tang F. AnchorCrafter: Animate cyber-anchors
selling your products via human-object interacting video generation. arXiv:2411.17383, 2024.
[48] Men YF, Yao Y, Cui MM, Bo LF. MIMO: Controllable character video synthesis with spatial decomposed modeling. In: Proc. of the
2025 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Nashville: IEEE, 2025. 21181–21191. [doi: 10.1109/CVPR52734.
2025.01973]
[49] Hu L, Wang GY, Shen Z, Gao X, Meng DC, Zhuo L, Zhang P, Zhang B, Bo LF. Animate Anyone 2: High-fidelity character image
animation with environment affordance. arXiv:2502.06145, 2025.
[50] Wang ZX, Yuan ZY, Wang XT, Li YW, Chen TS, Xia MH, Luo P, Shan Y. MotionCtrl: A unified and flexible motion controller for
video generation. In: Proc. of the 2024 ACM SIGGRAPH Conf. Denver: ACM, 2024. 114. [doi: 10.1145/3641519.3657518]
[51] He H, Xu YH, Guo YW, Wetzstein G, Dai B, Li HS, Yang CY. CameraCtrl: Enabling camera control for text-to-video generation.
arXiv:2404.02101, 2024.
[52] Shi XY, Huang ZY, Wang FY, Bian WK, Li DS, Zhang Y, Zhang MY, Cheung KC, See S, Qin HW, Dai JF, Li HS. Motion-I2V:
Consistent and controllable image-to-video generation with explicit motion modeling. In: Proc. of the 2024 ACM SIGGRAPH Conf.
Denver: ACM, 2024. 111. [doi: 10.1145/3641519.3657497]
[53] Niu MY, Cun XD, Wang XT, Zhang Y, Shan Y, Zheng YQ. MOFA-Video: Controllable image animation via generative motion field

