Page 335 - 《软件学报》2026年第4期
P. 335
1776 软件学报 2026 年第 37 卷第 4 期
the 8th Int’l Conf. on Learning Representations. Addis Ababa: ICLR, 2020.
[64] Sun C, Myers A, Vondrick C, Murphy K, Schmid C. VideoBERT: A joint model for video and language representation learning. In: Proc.
of the 2019 IEEE/CVF Int’l Conf. on Computer Vision. Seoul: IEEE, 2019. 7463–7472. [doi: 10.1109/ICCV.2019.00756]
[65] Han TD, Xie WD, Zisserman A. Temporal alignment networks for long-term video. In: Proc. of the 2022 IEEE/CVF Conf. on Computer
Vision and Pattern Recognition. New Orleans: IEEE, 2022. 2896–2906. [doi: 10.1109/CVPR52688.2022.00292]
作者简介
吴益露, 博士, 主要研究领域为教学视频理解, 技能学习.
王瀚霖, 硕士, 主要研究领域为深度学习, 视频理解, 视频生成.
王利民, 博士, 教授, 博士生导师, CCF 高级会员, 主要研究领域为计算机视觉, 深度学习, 视频理解.

