Page 82 - 《软件学报》2026年第5期
P. 82
方承炀 等: 面向免训练视频问答的双重自适应冗余消除 1961
实现了关键帧序列的自适应选择和误差修正, 既减少了计算冗余, 又保持了时间动态的完整性. 此外, 提出的动态
空间采样策略有效抑制了空间噪声干扰, 增强了目标区域的语义相关性. 我们在 4 个广泛使用的公开数据集上进
行了大量实验. 结果表明, 在定量分析中, 该方法在准确率和质量指标上显著优于 14 个最新比较模型. 可视化结果
展示了模型在复杂场景中 (如多人交互和细粒度动作) 依然能准确定位时空关键信息. 然而, 现有方法在处理长视
频时仍面临视频理解的瓶颈. 未来工作将探索时序跳跃采样策略, 以进一步提升模型对视频叙事结构的理解能力.
References
[1] Hu JX, Meng ZH. Research on video question answering based on video description and reading comprehension. Application Research of
Computers, 2021, 38(12): 3781–3785 (in Chinese with English abstract). [doi: 10.19734/j.issn.1001-3695.2021.04.0152]
[2] Yao X, Gao JY, Xu CS. Self-supervised graph contrastive learning for video question answering. Ruan Jian Xue Bao/Journal of Software,
2023, 34(5): 2083–2100 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/6775.htm [doi: 10.13328/j.cnki.jos.006775]
[3] Zhao EY, Song N, Nie J, Wang X, Zheng CY, Wei ZQ. Scale-guided fusion inference network for remote sensing visual question
answering. Ruan Jian Xue Bao/Journal of Software, 2024, 35(5): 2133–2149 (in Chinese with English abstract). http://www.jos.org.cn/
1000-9825/7025.htm [doi: 10.13328/j.cnki.jos.007025]
[4] Ying J, Zhang ZD, Gao YH, Yang ZW, Li L, Xiao M, Sun YQ, Yan CG. Survey on vision-language pre-training. Ruan Jian Xue
Bao/Journal of Software, 2023, 34(5): 2000–2023 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/6774.htm [doi: 10.
13328/j.cnki.jos.006774]
[5] Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, Krueger G, Sutskever I. Learning
transferable visual models from natural language supervision. In: Proc. of the 38th Int’l Conf. on Machine Learning. PMLR, 2021.
8748–8763.
[6] Kim W, Choi C, Lee W, Rhee W. An image grid can be worth a video: Zero-shot video question answering using a VLM. IEEE Access,
2024, 12: 193057–193075. [doi: 10.1109/ACCESS.2024.3517625]
[7] Wu WH. FreeVA: Offline MLLM as training-free video assistant. arXiv:2405.07798, 2024.
[8] Xu MZ, Gao MF, Gan Z, Chen HY, Lai ZF, Gang HM, Kang K, Dehghan A. SlowFast-LLaVA: A strong training-free baseline for video
large language models. arXiv:2407.15841, 2024.
[9] Han K, Guo JY, Tang YH, He W, Wu EH, Wang YH. Free Video-LLM: Prompt-guided visual perception for efficient training-free video
LLMs. arXiv:2410.10441, 2024.
[10] Chen D, Dolan W. Collecting highly parallel data for paraphrase evaluation. In: Proc. of the 49th Annual Meeting of the Association for
Computational Linguistics: Human Language Technologies. Portland: ACL, 2011. 190–200.
[11] Xu J, Mei T, Yao T, Rui Y. MSR-VTT: A large video description dataset for bridging video and language. In: Proc. of the 2016 IEEE
Conf. on Computer Vision and Pattern Recognition. Las Vegas: IEEE, 2016. 5288–5296. [doi: 10.1109/CVPR.2016.571]
[12] Jang Y, Song Y, Yu Y, Kim Y, Kim G. TGIF-QA: Toward spatio-temporal reasoning in visual question answering. In: Proc. of the 2017
IEEE Conf. on Computer Vision and Pattern Recognition. Honolulu: IEEE, 2017. 1359–1367. [doi: 10.1109/CVPR.2017.149]
[13] Heilbron FC, Escorcia V, Ghanem B, Niebles JC. ActivityNet: A large-scale video benchmark for human activity understanding. In: Proc.
of the 2015 IEEE Conf. on Computer Vision and Pattern Recognition. Boston: IEEE, 2015. 961–970. [doi: 10.1109/CVPR.2015.
7298698]
[14] Brown TB, Mann B, Ryder N, et al. Language models are few-shot learners. In: Proc. of the 34th Int’l Conf. on Neural Information
Processing Systems. Vancouver: Curran Associates Inc., 2020. 1877–1901.
[15] Touvron H, Lavril T, Izacard G, Martinet X, Lachaux MA, Lacroix T, Rozière B, Goyal N, Hambro E, Azhar F, Rodriguez A, Joulin A,
Grave E, Lample G. LLaMA: Open and efficient foundation language models. arXiv:2302.13971, 2023.
[16] Alayrac JB, Donahue J, Luc P, et al. Flamingo: A visual language model for few-shot learning. In: Proc. of the 36th Int’l Conf. on Neural
Information Processing Systems. New Orleans: Curran Associates Inc., 2022. 23716–23736.
[17] Li JN, Li DX, Savarese S, Hoi S. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language
models. In: Proc. of the 40th Int’l Conf. on Machine Learning. Honolulu: JMLR.org, 2023. 19730–19742.
[18] Liu HT, Li CY, Wu QY, Lee YJ. Visual instruction tuning. In: Proc. of the 37th Int’l Conf. on Neural Information Processing Systems.
New Orleans: Curran Associates Inc., 2023. 34892–34916.
[19] Liu HT, Li CY, Li YH, Li B, Zhang YH, Shen S, Lee YJ. LLaVA-NeXT: Improved reasoning, OCR, and world knowledge. 2024. https://
llava-vl.github.io/blog/2024-01-30-llava-next/
[20] Dai WL, Li JN, Li DX, Tiong AMH, Zhao JQ, Wang WS, Li BY, Fung P, Hoi S. InstructBLIP: Towards general-purpose vision-language

