Page 332 - 《软件学报》2026年第4期
P. 332
吴益露 等: 基于分类检索的操作规划方法 1773
从而满足不同下游任务的需求.
References
[1] Carreira J, Zisserman A. Quo vadis, action recognition? A new model and the kinetics dataset. In: Proc. of the 2017 IEEE Conf. on
Computer Vision and Pattern Recognition. Honolulu: IEEE, 2017. 4724–4733. [doi: 10.1109/CVPR.2017.502]
[2] Wang XL, Girshick R, Gupta A, He KM. Non-local neural networks. In: Proc. of the 2018 IEEE/CVF Conf. on Computer Vision and
Pattern Recognition. Salt Lake City: IEEE, 2018. 7794–7803. [doi: 10.1109/CVPR.2018.00813]
[3] Tong Z, Song YB, Wang J, Wang LM. VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-
training. In: Proc. of the 36th Conf. on Neural Information Processing Systems. New Orleans: Curran Associates Inc., 2022.
[4] Wang LM, Huang BK, Zhao ZY, Tong Z, He YN, Wang Y, Wang YL, Qiao Y. VideoMAE V2: Scaling video masked autoencoders with
dual masking. In: Proc. of the 2023 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023.
14549–14560. [doi: 10.1109/CVPR52729.2023.01398]
[5] Li KC, He YN, Wang Y, Li YZ, Wang WH, Luo P, Wang YL, Wang LM, Qiao Y. VideoChat: Chat-centric video understanding.
arXiv:2305.06355, 2023.
[6] Zhou BL, Andonian A, Oliva A, Torralba A. Temporal relational reasoning in videos. In: Proc. of the 15th European Conf. on Computer
Vision. Munich: Springer, 2018. 831–846. [doi: 10.1007/978-3-030-01246-5_49]
[7] Feichtenhofer C, Fan HQ, Malik J, He KM. SlowFast networks for video recognition. In: Proc. of the 2019 IEEE/CVF Int’l Conf. on
Computer Vision. Seoul: IEEE, 2019. 6201–6210. [doi: 10.1109/ICCV.2019.00630]
[8] Wang LM, Xiong YJ, Wang Z, Qiao Y, Lin DH, Tang XO, van Gool L. Temporal segment networks for action recognition in videos.
IEEE Trans. on Pattern Analysis and Machine Intelligence, 2019, 41(11): 2740–2755. [doi: 10.1109/TPAMI.2018.2868668]
[9] Lin TW, Liu X, Li X, Ding ER, Wen SL. BMN: Boundary-matching network for temporal action proposal generation. In: Proc. of the
2019 IEEE/CVF Int’l Conf. on Computer Vision. Seoul: IEEE, 2019. 3888–3897. [doi: 10.1109/ICCV.2019.00399]
[10] Lin TW, Zhao X, Su HS, Wang CJ, Yang M. BSN: Boundary sensitive network for temporal action proposal generation. In: Proc. of the
15th European Conf. on Computer Vision. Munich: Springer, 2018. 3–21. [doi: 10.1007/978-3-030-01225-0_1]
[11] Song L, Zhang SW, Yu G, Sun HB. TACNet: Transition-aware context network for spatio-temporal action detection. In: Proc. of the
2019 IEEE Conf. on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2019. 11979–11987. [doi: 10.1109/CVPR.2019.
01226]
[12] Yang XT, Yang XD, Liu MY, Xiao FY, Davis LS, Kautz J. STEP: Spatio-temporal progressive learning for video action detection. In:
Proc. of the 2019 IEEE Conf. on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2019. 264–272. [doi: 10.1109/CVPR.
2019.00035]
[13] Farha YA, Ke QH, Schiele B, Gall J. Long-term anticipation of activities with cycle consistency. In: Proc. of the 42nd German Conf. on
Pattern Recognition. Tübingen: Springer, 2020. 159–173. [doi: 10.1007/978-3-030-71278-5_12]
[14] Alayrac JB, Bojanowski P, Agrawal N, Sivic J, Laptev I, Lacoste-Julien S. Unsupervised learning from narrated instruction videos. In:
Proc. of the 2016 IEEE Conf. on Computer Vision and Pattern Recognition. Las Vegas: IEEE, 2016. 4575–4583. [doi: 10.1109/CVPR.
2016.495]
[15] Zhukov D, Alayrac JB, Cinbis RG, Fouhey D, Laptev I, Sivic J. Cross-task weakly supervised learning from instructional videos. In:
Proc. of the 2019 IEEE Conf. on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2019. 3532–3540. [doi: 10.1109/CVPR.
2019.00365]
[16] Miech A, Zhukov D, Alayrac JB, Tapaswi M, Laptev I, Sivic J. HowTo100M: Learning a text-video embedding by watching hundred
million narrated video clips. In: Proc. of the 2019 IEEE/CVF Int’l Conf. on Computer Vision. Seoul: IEEE, 2019. 2630–2640. [doi: 10.
1109/ICCV.2019.00272]
[17] Tang YS, Lu JW, Zhou J. Comprehensive instructional video analysis: The COIN dataset and performance evaluation. IEEE Trans. on
Pattern Analysis and Machine Intelligence, 2021, 43(9): 3138–3153. [doi: 10.1109/TPAMI.2020.2980824]
[18] Chang CY, Huang DA, Xu DF, Adeli E, Li FF, Niebles JC. Procedure planning in instructional videos. In: Proc. of the 16th European
Conf. on Computer Vision. Glasgow: Springer, 2020. 334–350. [doi: 10.1007/978-3-030-58621-8_20]
[19] Finn C, Tan XY, Duan Y, Darrell T, Levine S, Abbeel P. Deep spatial autoencoders for visuomotor learning. In: Proc. of the 2016 IEEE
Int’l Conf. on Robotics and Automation. Stockholm: IEEE, 2016. 512–519. [doi: 10.1109/ICRA.2016.7487173]
[20] Finn C, Levine S. Deep visual foresight for planning robot motion. In: Proc. of the 2017 IEEE Int’l Conf. on Robotics and Automation.
Singapore: IEEE, 2017. 2786–2793. [doi: 10.1109/ICRA.2017.7989324]
[21] Zhao H, Hadji I, Dvornik N, Derpanis KG, Wildes RP, Jepson AD. P³IV: Probabilistic procedure planning from instructional videos with

