Page 335 - 《软件学报》2026年第4期
P. 335

1776                                                       软件学报  2026  年第  37  卷第  4  期


                     the 8th Int’l Conf. on Learning Representations. Addis Ababa: ICLR, 2020.
                 [64]   Sun C, Myers A, Vondrick C, Murphy K, Schmid C. VideoBERT: A joint model for video and language representation learning. In: Proc.
                     of the 2019 IEEE/CVF Int’l Conf. on Computer Vision. Seoul: IEEE, 2019. 7463–7472. [doi: 10.1109/ICCV.2019.00756]
                 [65]   Han TD, Xie WD, Zisserman A. Temporal alignment networks for long-term video. In: Proc. of the 2022 IEEE/CVF Conf. on Computer
                     Vision and Pattern Recognition. New Orleans: IEEE, 2022. 2896–2906. [doi: 10.1109/CVPR52688.2022.00292]

                 作者简介
                 吴益露, 博士, 主要研究领域为教学视频理解, 技能学习.
                 王瀚霖, 硕士, 主要研究领域为深度学习, 视频理解, 视频生成.
                 王利民, 博士, 教授, 博士生导师, CCF  高级会员, 主要研究领域为计算机视觉, 深度学习, 视频理解.
   330   331   332   333   334   335   336   337   338   339   340