Page 443 - 《软件学报》2026年第4期
P. 443
1884 软件学报 2026 年第 37 卷第 4 期
arXiv:2401.09962, 2024.
[116] Raj A, Kaza S, Poole B, Niemeyer M, Ruiz N, Mildenhall B, Zada S, Aberman K, Rubinstein M, Barron J, Li YZ, Jampani V.
DreamBooth3D: Subject-driven text-to-3D generation. In: Proc. of the 2023 IEEE/CVF Int’l Conf. on Computer Vision. Paris: IEEE,
2023. 2349–2359. [doi: 10.1109/ICCV51070.2023.00223]
[117] Ouyang YC, Chai WH, Ye JY, Tao DP, Zhan YB, Wang GA. Chasing consistency in text-to-3D generation from a single image.
arXiv:2309.03599, 2023.
[118] Hsiao TF, Ruan BK, Wu YL, Lin TL, Shuai HH. TF-TI2I: Training-free text-and-image-to-image generation via multi-modal implicit-
context learning in text-to-image models. arXiv:2503.15283, 2025.
[119] Sung-Bin K, Senocak A, Ha H, Owens A, Oh TH. Sound to visual scene generation by audio-to-visual latent alignment. In: Proc. of the
2023 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023. 6430–6440. [doi: 10.1109/CVPR52729.
2023.00622]
[120] Petermann D, Kalayeh MM. Seeing sound: Assembling sounds from visuals for audio-to-image generation. arXiv:2501.05413, 2025.
[121] Lopez E, Sigillo L, Colonnese F, Panella M, Comminiello D. Guess what I think: Streamlined EEG-to-image generation with latent
diffusion models. In: Proc. of the 2025 IEEE Int’l Conf. on Acoustics, Speech and Signal Processing. Hyderabad: IEEE, 2025. 1–5. [doi:
10.1109/ICASSP49660.2025.10890059]
[122] Yang FY, Zhang JC, Owens A. Generating visual scenes from touch. In: Proc. of the 2023 IEEE/CVF Int’l Conf. on Computer Vision.
Paris: IEEE, 2023. 22013–22023. [doi: 10.1109/ICCV51070.2023.02017]
[123] Zhou CT, Yu LL, Babu A, Tirumala K, Yasunaga M, Shamis L, Kahn J, Ma XZ, Zettlemoyer L, Levy O. Transfusion: Predict the next
token and diffuse images with one multi-modal model. In: Proc. of the 13th Int’l Conf. on Learning Representations. Singapore:
OpenReview.net, 2025. 1–24.
[124] Xie JH, Mao WJ, Bai ZC, Zhang DJ, Wang WH, Lin KQ, Gu YC, Chen ZJ, Yang ZH, Shou MZ. Show-O: One single Transformer to
unify multimodal understanding and generation. In: Proc. of the 13th Int’l Conf. on Learning Representations. Singapore:
OpenReview.net, 2025. 1–25.
附中文参考文献
[11] 左然, 胡皓翔, 邓小明, 马翠霞, 王宏安. 基于手绘草图的视觉内容生成深度学习方法综述. 软件学报, 2024, 35(7): 3497–3530. http://
www.jos.org.cn/1000-9825/7053.htm [doi: 10.13328/j.cnki.jos.007053]
[12] 余凯, 宾燚, 郑自强, 杨阳. 基于条件语义增强的文本到图像生成. 软件学报, 2024, 35(5): 2150–2164. http://www.jos.org.cn/1000-
9825/7024.htm [doi: 10.13328/j.cnki.jos.007024]
作者简介
何子健, 博士, 主要研究领域为计算机视觉, 图像视频生成, 多模态融合.
李冠彬, 博士, 教授, 博士生导师, CCF 杰出会员, 主要研究领域为视觉内容理解与生成.

