Page 443 - 《软件学报》2026年第4期
P. 443

1884                                                       软件学报  2026  年第  37  卷第  4  期


                      arXiv:2401.09962, 2024.
                 [116]   Raj  A,  Kaza  S,  Poole  B,  Niemeyer  M,  Ruiz  N,  Mildenhall  B,  Zada  S,  Aberman  K,  Rubinstein  M,  Barron  J,  Li  YZ,  Jampani  V.
                      DreamBooth3D: Subject-driven text-to-3D generation. In: Proc. of the 2023 IEEE/CVF Int’l Conf. on Computer Vision. Paris: IEEE,
                      2023. 2349–2359. [doi: 10.1109/ICCV51070.2023.00223]
                 [117]   Ouyang  YC,  Chai  WH,  Ye  JY,  Tao  DP,  Zhan  YB,  Wang  GA.  Chasing  consistency  in  text-to-3D  generation  from  a  single  image.
                      arXiv:2309.03599, 2023.
                 [118]   Hsiao TF, Ruan BK, Wu YL, Lin TL, Shuai HH. TF-TI2I: Training-free text-and-image-to-image generation via multi-modal implicit-
                      context learning in text-to-image models. arXiv:2503.15283, 2025.
                 [119]   Sung-Bin K, Senocak A, Ha H, Owens A, Oh TH. Sound to visual scene generation by audio-to-visual latent alignment. In: Proc. of the
                      2023 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023. 6430–6440. [doi: 10.1109/CVPR52729.
                      2023.00622]
                 [120]   Petermann D, Kalayeh MM. Seeing sound: Assembling sounds from visuals for audio-to-image generation. arXiv:2501.05413, 2025.
                 [121]   Lopez E, Sigillo L, Colonnese F, Panella M, Comminiello D. Guess what I think: Streamlined EEG-to-image generation with latent
                      diffusion models. In: Proc. of the 2025 IEEE Int’l Conf. on Acoustics, Speech and Signal Processing. Hyderabad: IEEE, 2025. 1–5. [doi:
                      10.1109/ICASSP49660.2025.10890059]
                 [122]   Yang FY, Zhang JC, Owens A. Generating visual scenes from touch. In: Proc. of the 2023 IEEE/CVF Int’l Conf. on Computer Vision.
                      Paris: IEEE, 2023. 22013–22023. [doi: 10.1109/ICCV51070.2023.02017]
                 [123]   Zhou CT, Yu LL, Babu A, Tirumala K, Yasunaga M, Shamis L, Kahn J, Ma XZ, Zettlemoyer L, Levy O. Transfusion: Predict the next
                      token  and  diffuse  images  with  one  multi-modal  model.  In:  Proc.  of  the  13th  Int’l  Conf.  on  Learning  Representations.  Singapore:
                      OpenReview.net, 2025. 1–24.
                 [124]   Xie JH, Mao WJ, Bai ZC, Zhang DJ, Wang WH, Lin KQ, Gu YC, Chen ZJ, Yang ZH, Shou MZ. Show-O: One single Transformer to
                      unify  multimodal  understanding  and  generation.  In:  Proc.  of  the  13th  Int’l  Conf.  on  Learning  Representations.  Singapore:
                      OpenReview.net, 2025. 1–25.

                 附中文参考文献
                 [11]   左然, 胡皓翔, 邓小明, 马翠霞, 王宏安. 基于手绘草图的视觉内容生成深度学习方法综述. 软件学报, 2024, 35(7): 3497–3530. http://
                     www.jos.org.cn/1000-9825/7053.htm [doi: 10.13328/j.cnki.jos.007053]
                 [12]   余凯, 宾燚, 郑自强, 杨阳. 基于条件语义增强的文本到图像生成. 软件学报, 2024, 35(5): 2150–2164. http://www.jos.org.cn/1000-
                     9825/7024.htm [doi: 10.13328/j.cnki.jos.007024]

                 作者简介
                 何子健, 博士, 主要研究领域为计算机视觉, 图像视频生成, 多模态融合.
                 李冠彬, 博士, 教授, 博士生导师, CCF  杰出会员, 主要研究领域为视觉内容理解与生成.
   438   439   440   441   442   443   444