Page 83 - 《软件学报》2026年第5期
P. 83

1962                                                       软件学报  2026  年第  37  卷第  5  期


                     models  with  instruction  tuning.  In:  Proc.  of  the  37th  Int’l  Conf.  on  Neural  Information  Processing  Systems.  New  Orleans:  Curran
                     Associates Inc., 2023. 49250–49267.
                 [21]   Zhu DY, Chen J, Shen XQ, Li X, Elhoseiny M. MiniGPT-4: Enhancing vision-language understanding with advanced large language
                     models. arXiv:2304.10592, 2023.
                 [22]   Hong YN, Zhen HY, Chen PH, Zheng SH, Du YL, Chen ZF, Gan C. 3D-LLM: Injecting the 3D world into large language models. In:
                     Proc. of the 37th Int’l Conf. on Neural Information Processing Systems. New Orleans: Curran Associates Inc., 2023. 20482–20494.
                 [23]   McKinzie B, Gan Z, Fauconnier JP, et al. MM1: Methods, analysis & insights from multimodal LLM pre-training. arXiv:2403.09611,
                     2024.
                 [24]   You HX, Zhang HT, Gan Z, Du XZ, Zhang BW, Wang ZR, Cao LL, Chang SF, Yang YF. Ferret: Refer and ground anything anywhere at
                     any granularity. arXiv:2310.07704, 2023.
                 [25]   Zhang HT, You HX, Dufter P, Zhang BW, Chen C, Chen HY, Fu TJ, Wang WY, Chang SF, Gan Z, Yang YF. Ferret-v2: An improved
                     baseline for referring and grounding with large language models. arXiv:2404.07973, 2024.
                 [26]   Bai  JZ,  Bai  S,  Yang  SS,  Wang  SJ,  Tan  SN,  Wang  P,  Lin  JY,  Zhou  C,  Zhou  JR.  Qwen-VL:  A  versatile  vision-language  model  for
                     understanding, localization, text reading, and beyond. arXiv:2308.12966, 2023.
                 [27]   Yang A, Miech A, Sivic J, Laptev I, Schmid C. Zero-shot video question answering via frozen bidirectional language models. In: Proc. of
                     the 36th Int’l Conf. on Neural Information Processing Systems. New Orleans: Curran Associates Inc., 2022. 124–141.
                 [28]   Li KC, He Y, Wang Y, Li YZ, Wang WH, Luo P, Wang YL, Wang LM, Qiao Y. VideoChat: Chat-centric video understanding. arXiv:
                     2305.06355, 2024.
                 [29]   Zhang H, Li X, Bing LD. Video-LLaMA: An instruction-tuned audio-visual language model for video understanding. In: Proc. of the
                     2023 Conf. on Empirical Methods in Natural Language Processing: System Demonstrations. Singapore: ACL, 2023. 543–553. [doi: 10.
                     18653/v1/2023.emnlp-demo.49]
                 [30]   Maaz M, Rasheed H, Khan S, Khan F. Video-ChatGPT: Towards detailed video understanding via large vision and language models. In:
                     Proc.  of  the  62nd  Annual  Meeting  of  the  Association  for  Computational  Linguistics  (Vol.  1:  Long  Papers).  Bangkok:  ACL,  2024.
                     12585–12602. [doi: 10.18653/v1/2024.acl-long.679]
                 [31]   Jin P, Takanobu R, Zhang WC, Cao XC, Yuan L. Chat-UniVi: Unified visual representation empowers large language models with image
                     and  video  understanding.  In:  Proc.  of  the  2024  IEEE/CVF  Conf.  on  Computer  Vision  and  Pattern  Recognition.  Seattle:  IEEE,  2024.
                     13700–13710. [doi: 10.1109/CVPR52733.2024.01300]
                 [32]   Xu L, Zhao YL, Zhou DQ, Lin ZJ, Ng SK, Feng JS. PLLaVA: Parameter-free LLaVA extension from images to videos for video dense
                     captioning. arXiv:2404.16994, 2024.
                 [33]   Zhang YH, Li B, Liu HT, Lee YJ, Gui LK, Fu D, Feng JS, Liu ZW, Li CY. LLaVA-NeXT: A strong zero-shot video understanding
                     model. 2024. https://llava-vl.github.io/blog/2024-04-30-llava-next-video/
                 [34]   Ren SH, Yao LL, Li SC, Sun X, Hou L. TimeChat: A time-sensitive multimodal large language model for long video understanding. In:
                     Proc. of the 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024. 14313–14323. [doi: 10.1109/
                     CVPR52733.2024.01357]
                 [35]   Li YW, Wang CY, Jia JY. LLaMA-VID: An image is worth 2 tokens in large language models. In: Proc. of the 18th European Conf. on
                     Computer Vision. Milan: Springer, 2025. 323–340. [doi: 10.1007/978-3-031-72952-2_19]
                 [36]   Liu ZY, Dong YH, Liu ZW, Hu W, Lu JW, Rao YM. Oryx MLLM: On-demand spatial-temporal understanding at arbitrary resolution.
                     arXiv:2409.12961, 2025.
                 [37]   Shang YZ, Xu BX, Kang WT, Cai M, Li YH, Wen ZH, Dong Z, Keutzer K, Lee YJ, Yan Y. Interpolating Video-LLMs: Toward longer-
                     sequence LMMs in a training-free manner. arXiv:2409.12963, 2024.
                 [38]   Guo D, Yao ST, Wang H, Wang M. Embedding VLAD in Transformer for video question answering. Chinese Journal of Computers,
                     2023, 46(4): 671–689 (in Chinese with English abstract). [doi: 10.11897/SP.J.1016.2023.00671]
                 [39]   Lin  B,  Ye  Y,  Zhu  B,  Cui  JX,  Ning  MN,  Jin  P,  Yuan  L.  Video-LLaVA:  Learning  united  visual  representation  by  alignment  before
                     projection. In: Proc. of the 2024 Conf. on Empirical Methods in Natural Language Processing. Miami: ACL, 2024. 5971–5984. [doi: 10.
                     18653/v1/2024.emnlp-main.342]
                 [40]   Ma F, Jin X, Wang H, Yang J, Luo P. Vista-LLaMA: Reducing hallucination in video narrator via equal distance to visual tokens. In:
                     Proc. of the 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). Seattle: IEEE, 2024. 13151–13160. [doi: 10.
                     1109/CVPR52733.2024.01249]
                 [41]   Li KC, Wang YL, He Y, Li YZ, Wang Y, Liu Y, Wang Z, Xu JL, Chen G, Luo P, Wang LM, Qiao Y. MVBench: A comprehensive multi-
                     modal video understanding benchmark. In: Proc. of the 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Seattle:
   78   79   80   81   82   83   84   85   86   87   88