Page 83 - 《软件学报》2026年第5期
P. 83
1962 软件学报 2026 年第 37 卷第 5 期
models with instruction tuning. In: Proc. of the 37th Int’l Conf. on Neural Information Processing Systems. New Orleans: Curran
Associates Inc., 2023. 49250–49267.
[21] Zhu DY, Chen J, Shen XQ, Li X, Elhoseiny M. MiniGPT-4: Enhancing vision-language understanding with advanced large language
models. arXiv:2304.10592, 2023.
[22] Hong YN, Zhen HY, Chen PH, Zheng SH, Du YL, Chen ZF, Gan C. 3D-LLM: Injecting the 3D world into large language models. In:
Proc. of the 37th Int’l Conf. on Neural Information Processing Systems. New Orleans: Curran Associates Inc., 2023. 20482–20494.
[23] McKinzie B, Gan Z, Fauconnier JP, et al. MM1: Methods, analysis & insights from multimodal LLM pre-training. arXiv:2403.09611,
2024.
[24] You HX, Zhang HT, Gan Z, Du XZ, Zhang BW, Wang ZR, Cao LL, Chang SF, Yang YF. Ferret: Refer and ground anything anywhere at
any granularity. arXiv:2310.07704, 2023.
[25] Zhang HT, You HX, Dufter P, Zhang BW, Chen C, Chen HY, Fu TJ, Wang WY, Chang SF, Gan Z, Yang YF. Ferret-v2: An improved
baseline for referring and grounding with large language models. arXiv:2404.07973, 2024.
[26] Bai JZ, Bai S, Yang SS, Wang SJ, Tan SN, Wang P, Lin JY, Zhou C, Zhou JR. Qwen-VL: A versatile vision-language model for
understanding, localization, text reading, and beyond. arXiv:2308.12966, 2023.
[27] Yang A, Miech A, Sivic J, Laptev I, Schmid C. Zero-shot video question answering via frozen bidirectional language models. In: Proc. of
the 36th Int’l Conf. on Neural Information Processing Systems. New Orleans: Curran Associates Inc., 2022. 124–141.
[28] Li KC, He Y, Wang Y, Li YZ, Wang WH, Luo P, Wang YL, Wang LM, Qiao Y. VideoChat: Chat-centric video understanding. arXiv:
2305.06355, 2024.
[29] Zhang H, Li X, Bing LD. Video-LLaMA: An instruction-tuned audio-visual language model for video understanding. In: Proc. of the
2023 Conf. on Empirical Methods in Natural Language Processing: System Demonstrations. Singapore: ACL, 2023. 543–553. [doi: 10.
18653/v1/2023.emnlp-demo.49]
[30] Maaz M, Rasheed H, Khan S, Khan F. Video-ChatGPT: Towards detailed video understanding via large vision and language models. In:
Proc. of the 62nd Annual Meeting of the Association for Computational Linguistics (Vol. 1: Long Papers). Bangkok: ACL, 2024.
12585–12602. [doi: 10.18653/v1/2024.acl-long.679]
[31] Jin P, Takanobu R, Zhang WC, Cao XC, Yuan L. Chat-UniVi: Unified visual representation empowers large language models with image
and video understanding. In: Proc. of the 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024.
13700–13710. [doi: 10.1109/CVPR52733.2024.01300]
[32] Xu L, Zhao YL, Zhou DQ, Lin ZJ, Ng SK, Feng JS. PLLaVA: Parameter-free LLaVA extension from images to videos for video dense
captioning. arXiv:2404.16994, 2024.
[33] Zhang YH, Li B, Liu HT, Lee YJ, Gui LK, Fu D, Feng JS, Liu ZW, Li CY. LLaVA-NeXT: A strong zero-shot video understanding
model. 2024. https://llava-vl.github.io/blog/2024-04-30-llava-next-video/
[34] Ren SH, Yao LL, Li SC, Sun X, Hou L. TimeChat: A time-sensitive multimodal large language model for long video understanding. In:
Proc. of the 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024. 14313–14323. [doi: 10.1109/
CVPR52733.2024.01357]
[35] Li YW, Wang CY, Jia JY. LLaMA-VID: An image is worth 2 tokens in large language models. In: Proc. of the 18th European Conf. on
Computer Vision. Milan: Springer, 2025. 323–340. [doi: 10.1007/978-3-031-72952-2_19]
[36] Liu ZY, Dong YH, Liu ZW, Hu W, Lu JW, Rao YM. Oryx MLLM: On-demand spatial-temporal understanding at arbitrary resolution.
arXiv:2409.12961, 2025.
[37] Shang YZ, Xu BX, Kang WT, Cai M, Li YH, Wen ZH, Dong Z, Keutzer K, Lee YJ, Yan Y. Interpolating Video-LLMs: Toward longer-
sequence LMMs in a training-free manner. arXiv:2409.12963, 2024.
[38] Guo D, Yao ST, Wang H, Wang M. Embedding VLAD in Transformer for video question answering. Chinese Journal of Computers,
2023, 46(4): 671–689 (in Chinese with English abstract). [doi: 10.11897/SP.J.1016.2023.00671]
[39] Lin B, Ye Y, Zhu B, Cui JX, Ning MN, Jin P, Yuan L. Video-LLaVA: Learning united visual representation by alignment before
projection. In: Proc. of the 2024 Conf. on Empirical Methods in Natural Language Processing. Miami: ACL, 2024. 5971–5984. [doi: 10.
18653/v1/2024.emnlp-main.342]
[40] Ma F, Jin X, Wang H, Yang J, Luo P. Vista-LLaMA: Reducing hallucination in video narrator via equal distance to visual tokens. In:
Proc. of the 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). Seattle: IEEE, 2024. 13151–13160. [doi: 10.
1109/CVPR52733.2024.01249]
[41] Li KC, Wang YL, He Y, Li YZ, Wang Y, Liu Y, Wang Z, Xu JL, Chen G, Luo P, Wang LM, Qiao Y. MVBench: A comprehensive multi-
modal video understanding benchmark. In: Proc. of the 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Seattle:

