Page 82 - 《软件学报》2026年第5期
P. 82

方承炀 等: 面向免训练视频问答的双重自适应冗余消除                                                      1961


                 实现了关键帧序列的自适应选择和误差修正, 既减少了计算冗余, 又保持了时间动态的完整性. 此外, 提出的动态
                 空间采样策略有效抑制了空间噪声干扰, 增强了目标区域的语义相关性. 我们在                         4  个广泛使用的公开数据集上进
                 行了大量实验. 结果表明, 在定量分析中, 该方法在准确率和质量指标上显著优于                       14  个最新比较模型. 可视化结果
                 展示了模型在复杂场景中         (如多人交互和细粒度动作) 依然能准确定位时空关键信息. 然而, 现有方法在处理长视
                 频时仍面临视频理解的瓶颈. 未来工作将探索时序跳跃采样策略, 以进一步提升模型对视频叙事结构的理解能力.


                 References
                  [1]   Hu JX, Meng ZH. Research on video question answering based on video description and reading comprehension. Application Research of
                     Computers, 2021, 38(12): 3781–3785 (in Chinese with English abstract). [doi: 10.19734/j.issn.1001-3695.2021.04.0152]
                  [2]   Yao X, Gao JY, Xu CS. Self-supervised graph contrastive learning for video question answering. Ruan Jian Xue Bao/Journal of Software,
                     2023, 34(5): 2083–2100 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/6775.htm [doi: 10.13328/j.cnki.jos.006775]
                  [3]   Zhao  EY,  Song  N,  Nie  J,  Wang  X,  Zheng  CY,  Wei  ZQ.  Scale-guided  fusion  inference  network  for  remote  sensing  visual  question
                     answering. Ruan Jian Xue Bao/Journal of Software, 2024, 35(5): 2133–2149 (in Chinese with English abstract). http://www.jos.org.cn/
                     1000-9825/7025.htm [doi: 10.13328/j.cnki.jos.007025]
                  [4]   Ying  J,  Zhang  ZD,  Gao  YH,  Yang  ZW,  Li  L,  Xiao  M,  Sun  YQ,  Yan  CG.  Survey  on  vision-language  pre-training.  Ruan  Jian  Xue
                     Bao/Journal of Software, 2023, 34(5): 2000–2023 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/6774.htm [doi: 10.
                     13328/j.cnki.jos.006774]
                  [5]   Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, Krueger G, Sutskever I. Learning
                     transferable  visual  models  from  natural  language  supervision.  In:  Proc.  of  the  38th  Int’l  Conf.  on  Machine  Learning.  PMLR,  2021.
                     8748–8763.
                  [6]   Kim W, Choi C, Lee W, Rhee W. An image grid can be worth a video: Zero-shot video question answering using a VLM. IEEE Access,
                     2024, 12: 193057–193075. [doi: 10.1109/ACCESS.2024.3517625]
                  [7]   Wu WH. FreeVA: Offline MLLM as training-free video assistant. arXiv:2405.07798, 2024.
                  [8]   Xu MZ, Gao MF, Gan Z, Chen HY, Lai ZF, Gang HM, Kang K, Dehghan A. SlowFast-LLaVA: A strong training-free baseline for video
                     large language models. arXiv:2407.15841, 2024.
                  [9]   Han K, Guo JY, Tang YH, He W, Wu EH, Wang YH. Free Video-LLM: Prompt-guided visual perception for efficient training-free video
                     LLMs. arXiv:2410.10441, 2024.
                 [10]   Chen D, Dolan W. Collecting highly parallel data for paraphrase evaluation. In: Proc. of the 49th Annual Meeting of the Association for
                     Computational Linguistics: Human Language Technologies. Portland: ACL, 2011. 190–200.
                 [11]   Xu J, Mei T, Yao T, Rui Y. MSR-VTT: A large video description dataset for bridging video and language. In: Proc. of the 2016 IEEE
                     Conf. on Computer Vision and Pattern Recognition. Las Vegas: IEEE, 2016. 5288–5296. [doi: 10.1109/CVPR.2016.571]
                 [12]   Jang Y, Song Y, Yu Y, Kim Y, Kim G. TGIF-QA: Toward spatio-temporal reasoning in visual question answering. In: Proc. of the 2017
                     IEEE Conf. on Computer Vision and Pattern Recognition. Honolulu: IEEE, 2017. 1359–1367. [doi: 10.1109/CVPR.2017.149]
                 [13]   Heilbron FC, Escorcia V, Ghanem B, Niebles JC. ActivityNet: A large-scale video benchmark for human activity understanding. In: Proc.
                     of  the  2015  IEEE  Conf.  on  Computer  Vision  and  Pattern  Recognition.  Boston:  IEEE,  2015.  961–970.  [doi:  10.1109/CVPR.2015.
                     7298698]
                 [14]   Brown TB, Mann B, Ryder N, et al. Language models are few-shot learners. In: Proc. of the 34th Int’l Conf. on Neural Information
                     Processing Systems. Vancouver: Curran Associates Inc., 2020. 1877–1901.
                 [15]   Touvron H, Lavril T, Izacard G, Martinet X, Lachaux MA, Lacroix T, Rozière B, Goyal N, Hambro E, Azhar F, Rodriguez A, Joulin A,
                     Grave E, Lample G. LLaMA: Open and efficient foundation language models. arXiv:2302.13971, 2023.
                 [16]   Alayrac JB, Donahue J, Luc P, et al. Flamingo: A visual language model for few-shot learning. In: Proc. of the 36th Int’l Conf. on Neural
                     Information Processing Systems. New Orleans: Curran Associates Inc., 2022. 23716–23736.
                 [17]   Li JN, Li DX, Savarese S, Hoi S. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language
                     models. In: Proc. of the 40th Int’l Conf. on Machine Learning. Honolulu: JMLR.org, 2023. 19730–19742.
                 [18]   Liu HT, Li CY, Wu QY, Lee YJ. Visual instruction tuning. In: Proc. of the 37th Int’l Conf. on Neural Information Processing Systems.
                     New Orleans: Curran Associates Inc., 2023. 34892–34916.
                 [19]   Liu HT, Li CY, Li YH, Li B, Zhang YH, Shen S, Lee YJ. LLaVA-NeXT: Improved reasoning, OCR, and world knowledge. 2024. https://
                     llava-vl.github.io/blog/2024-01-30-llava-next/
                 [20]   Dai WL, Li JN, Li DX, Tiong AMH, Zhao JQ, Wang WS, Li BY, Fung P, Hoi S. InstructBLIP: Towards general-purpose vision-language
   77   78   79   80   81   82   83   84   85   86   87