Page 439 - 《软件学报》2026年第4期
P. 439

1880                                                       软件学报  2026  年第  37  卷第  4  期


                 [32]   Wei ZC, Su QK, Qin L, Wang WZ. MM-Diff: High-fidelity image personalization via multi-modal condition integration. arXiv:2403.
                      15059, 2024.
                 [33]   He JJ, Geng YF, Bo LF. UniPortrait: A unified framework for identity-preserving single-and multi-human image personalization. arXiv:
                      2408.05939, 2024.
                 [34]   Xiao GX, Yin TW, Freeman WT, Durand F, Han S. FastComposer: Tuning-free multi-subject image generation with localized attention.
                      Int’l Journal of Computer Vision, 2025, 133(3): 1175–1194. [doi: 10.1007/s11263-024-02227-z]
                 [35]   Zhou  ZG,  Li  J,  Li  HX,  Chen  N,  Tang  X.  StoryMaker:  Towards  holistic  consistent  characters  in  text-to-image  generation.  arXiv:
                      2409.12576, 2024.
                 [36]   Zhao J, Zheng HL, Wang CY, Lan L, Huang WR, Tang YH. MagicNaming: Consistent identity generation by finding a “name space” in
                      T2I diffusion models. In: Proc. of the 39th AAAI Conf. on Artificial Intelligence. Philadelphia: AAAI Press, 2024. 10439–10447. [doi:
                      10.1609/aaai.v39i10.33133]
                 [37]   Nam J, Son S, Xu Z, Shi J, Liu DF, Liu F, Kim K, Zhou Y. Visual Persona: Foundation model for full-body human customization. In:
                      Proc. of the 2025 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Nashville: IEEE, 2025. 18630–18641. [doi: 10.1109/
                      CVPR52734.2025.01736]
                 [38]   Li H. Animate Anyone: Consistent and controllable image-to-video synthesis for character animation. In: Proc. of the 2024 IEEE/CVF
                      Conf. on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024. 8153–8163. [doi: 10.1109/CVPR52733.2024.00779]
                 [39]   Zhang Y, Gu JX, Wang LW, Wang H, Cheng JQ, Zhu YF, Zou FY. MimicMotion: High-quality human motion video generation with
                      confidence-aware pose guidance. arXiv:2406.19680, 2024.
                 [40]   Wang X, Zhang SW, Gao CX, Wang JY, Zhou XQ, Zhang YY, Yan LX, Sang N. UniAnimate: Taming unified video diffusion models
                      for consistent human image animation. arXiv:2406.01188, 2024.
                 [41]   Wang T, Li LJ, Lin K, Zhai YH, Lin CC, Yang ZY, Zhang HW, Liu ZC, Wang LJ. DisCo: Disentangled control for realistic human
                      dance generation. In: Proc. of the 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024. 9326–9336.
                      [doi: 10.1109/CVPR52733.2024.00891]
                 [42]   Xu ZC, Zhang JF, Liew JH, Yan HS, Liu JW, Zhang CX, Feng JS, Shou MZ. MagicAnimate: Temporally consistent human image
                      animation using diffusion model. In: Proc. of the 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Seattle: IEEE,
                      2024. 1481–1490. [doi: 10.1109/CVPR52733.2024.00147]
                 [43]   Zhu SH, Chen JL, Dai ZZ, Dong ZL, Xu YH, Cao X, Yao Y, Zhu H, Zhu SY. Champ: Controllable and consistent human image
                      animation with 3D parametric guidance. In: Proc. of the 18th European Conf. on Computer Vision. Milan: Springer, 2024. 145–162.
                      [doi: 10.1007/978-3-031-73001-6_9]
                 [44]   Parihar R, Gupta H, Vs S, Babu RV. Text2Place: Affordance-aware text guided human placement. In: Proc. of the 18th European Conf.
                      on Computer Vision. Milan: Springer, 2024. 57–77. [doi: 10.1007/978-3-031-72646-0_4]
                 [45]   Saini N, Bodla N, Shrivastava A, Ravichandran A, Zhang X, Shrivastava A, Singh B. InVi: Object insertion in videos using off-the-shelf
                      diffusion models. arXiv:2407.10958, 2024.
                 [46]   Qiu D, Chen ZY, Wang R, Fan MY, Yu CQ, Huang JS, Wen X. MovieCharacter: A tuning-free framework for controllable character
                      video synthesis. arXiv:2410.20974, 2024.
                 [47]   Xu ZY, Huang ZY, Cao J, Zhang Y, Cun XD, Shuai Q, Wang YC, Bao LC, Li JT, Tang F. AnchorCrafter: Animate cyber-anchors
                      selling your products via human-object interacting video generation. arXiv:2411.17383, 2024.
                 [48]   Men YF, Yao Y, Cui MM, Bo LF. MIMO: Controllable character video synthesis with spatial decomposed modeling. In: Proc. of the
                      2025 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Nashville: IEEE, 2025. 21181–21191. [doi: 10.1109/CVPR52734.
                      2025.01973]
                 [49]   Hu L, Wang GY, Shen Z, Gao X, Meng DC, Zhuo L, Zhang P, Zhang B, Bo LF. Animate Anyone 2: High-fidelity character image
                      animation with environment affordance. arXiv:2502.06145, 2025.
                 [50]   Wang ZX, Yuan ZY, Wang XT, Li YW, Chen TS, Xia MH, Luo P, Shan Y. MotionCtrl: A unified and flexible motion controller for
                      video generation. In: Proc. of the 2024 ACM SIGGRAPH Conf. Denver: ACM, 2024. 114. [doi: 10.1145/3641519.3657518]
                 [51]   He H, Xu YH, Guo YW, Wetzstein G, Dai B, Li HS, Yang CY. CameraCtrl: Enabling camera control for text-to-video generation.
                      arXiv:2404.02101, 2024.
                 [52]   Shi XY, Huang ZY, Wang FY, Bian WK, Li DS, Zhang Y, Zhang MY, Cheung KC, See S, Qin HW, Dai JF, Li HS. Motion-I2V:
                      Consistent and controllable image-to-video generation with explicit motion modeling. In: Proc. of the 2024 ACM SIGGRAPH Conf.
                      Denver: ACM, 2024. 111. [doi: 10.1145/3641519.3657497]
                 [53]   Niu MY, Cun XD, Wang XT, Zhang Y, Shan Y, Zheng YQ. MOFA-Video: Controllable image animation via generative motion field
   434   435   436   437   438   439   440   441   442   443   444