Page 332 - 《软件学报》2026年第4期
P. 332

吴益露 等: 基于分类检索的操作规划方法                                                            1773


                 从而满足不同下游任务的需求.

                 References
                  [1]   Carreira J, Zisserman A. Quo vadis, action recognition? A new model and the kinetics dataset. In: Proc. of the 2017 IEEE Conf. on
                     Computer Vision and Pattern Recognition. Honolulu: IEEE, 2017. 4724–4733. [doi: 10.1109/CVPR.2017.502]
                  [2]   Wang XL, Girshick R, Gupta A, He KM. Non-local neural networks. In: Proc. of the 2018 IEEE/CVF Conf. on Computer Vision and
                     Pattern Recognition. Salt Lake City: IEEE, 2018. 7794–7803. [doi: 10.1109/CVPR.2018.00813]
                  [3]   Tong  Z,  Song  YB,  Wang  J,  Wang  LM.  VideoMAE:  Masked  autoencoders  are  data-efficient  learners  for  self-supervised  video  pre-
                     training. In: Proc. of the 36th Conf. on Neural Information Processing Systems. New Orleans: Curran Associates Inc., 2022.
                  [4]   Wang LM, Huang BK, Zhao ZY, Tong Z, He YN, Wang Y, Wang YL, Qiao Y. VideoMAE V2: Scaling video masked autoencoders with
                     dual  masking.  In:  Proc.  of  the  2023  IEEE/CVF  Conf.  on  Computer  Vision  and  Pattern  Recognition.  Vancouver:  IEEE,  2023.
                     14549–14560. [doi: 10.1109/CVPR52729.2023.01398]
                  [5]   Li  KC,  He  YN,  Wang  Y,  Li  YZ,  Wang  WH,  Luo  P,  Wang  YL,  Wang  LM,  Qiao  Y.  VideoChat:  Chat-centric  video  understanding.
                     arXiv:2305.06355, 2023.
                  [6]   Zhou BL, Andonian A, Oliva A, Torralba A. Temporal relational reasoning in videos. In: Proc. of the 15th European Conf. on Computer
                     Vision. Munich: Springer, 2018. 831–846. [doi: 10.1007/978-3-030-01246-5_49]
                  [7]   Feichtenhofer C, Fan HQ, Malik J, He KM. SlowFast networks for video recognition. In: Proc. of the 2019 IEEE/CVF Int’l Conf. on
                     Computer Vision. Seoul: IEEE, 2019. 6201–6210. [doi: 10.1109/ICCV.2019.00630]
                  [8]   Wang LM, Xiong YJ, Wang Z, Qiao Y, Lin DH, Tang XO, van Gool L. Temporal segment networks for action recognition in videos.
                     IEEE Trans. on Pattern Analysis and Machine Intelligence, 2019, 41(11): 2740–2755. [doi: 10.1109/TPAMI.2018.2868668]
                  [9]   Lin TW, Liu X, Li X, Ding ER, Wen SL. BMN: Boundary-matching network for temporal action proposal generation. In: Proc. of the
                     2019 IEEE/CVF Int’l Conf. on Computer Vision. Seoul: IEEE, 2019. 3888–3897. [doi: 10.1109/ICCV.2019.00399]
                 [10]   Lin TW, Zhao X, Su HS, Wang CJ, Yang M. BSN: Boundary sensitive network for temporal action proposal generation. In: Proc. of the
                     15th European Conf. on Computer Vision. Munich: Springer, 2018. 3–21. [doi: 10.1007/978-3-030-01225-0_1]
                 [11]   Song L, Zhang SW, Yu G, Sun HB. TACNet: Transition-aware context network for spatio-temporal action detection. In: Proc. of the
                     2019  IEEE  Conf.  on  Computer  Vision  and  Pattern  Recognition.  Long  Beach:  IEEE,  2019.  11979–11987.  [doi:  10.1109/CVPR.2019.
                     01226]
                 [12]   Yang XT, Yang XD, Liu MY, Xiao FY, Davis LS, Kautz J. STEP: Spatio-temporal progressive learning for video action detection. In:
                     Proc. of the 2019 IEEE Conf. on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2019. 264–272. [doi: 10.1109/CVPR.
                     2019.00035]
                 [13]   Farha YA, Ke QH, Schiele B, Gall J. Long-term anticipation of activities with cycle consistency. In: Proc. of the 42nd German Conf. on
                     Pattern Recognition. Tübingen: Springer, 2020. 159–173. [doi: 10.1007/978-3-030-71278-5_12]
                 [14]   Alayrac JB, Bojanowski P, Agrawal N, Sivic J, Laptev I, Lacoste-Julien S. Unsupervised learning from narrated instruction videos. In:
                     Proc. of the 2016 IEEE Conf. on Computer Vision and Pattern Recognition. Las Vegas: IEEE, 2016. 4575–4583. [doi: 10.1109/CVPR.
                     2016.495]
                 [15]   Zhukov D, Alayrac JB, Cinbis RG, Fouhey D, Laptev I, Sivic J. Cross-task weakly supervised learning from instructional videos. In:
                     Proc. of the 2019 IEEE Conf. on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2019. 3532–3540. [doi: 10.1109/CVPR.
                     2019.00365]
                 [16]   Miech A, Zhukov D, Alayrac JB, Tapaswi M, Laptev I, Sivic J. HowTo100M: Learning a text-video embedding by watching hundred
                     million narrated video clips. In: Proc. of the 2019 IEEE/CVF Int’l Conf. on Computer Vision. Seoul: IEEE, 2019. 2630–2640. [doi: 10.
                     1109/ICCV.2019.00272]
                 [17]   Tang YS, Lu JW, Zhou J. Comprehensive instructional video analysis: The COIN dataset and performance evaluation. IEEE Trans. on
                     Pattern Analysis and Machine Intelligence, 2021, 43(9): 3138–3153. [doi: 10.1109/TPAMI.2020.2980824]
                 [18]   Chang CY, Huang DA, Xu DF, Adeli E, Li FF, Niebles JC. Procedure planning in instructional videos. In: Proc. of the 16th European
                     Conf. on Computer Vision. Glasgow: Springer, 2020. 334–350. [doi: 10.1007/978-3-030-58621-8_20]
                 [19]   Finn C, Tan XY, Duan Y, Darrell T, Levine S, Abbeel P. Deep spatial autoencoders for visuomotor learning. In: Proc. of the 2016 IEEE
                     Int’l Conf. on Robotics and Automation. Stockholm: IEEE, 2016. 512–519. [doi: 10.1109/ICRA.2016.7487173]
                 [20]   Finn C, Levine S. Deep visual foresight for planning robot motion. In: Proc. of the 2017 IEEE Int’l Conf. on Robotics and Automation.
                     Singapore: IEEE, 2017. 2786–2793. [doi: 10.1109/ICRA.2017.7989324]
                 [21]   Zhao H, Hadji I, Dvornik N, Derpanis KG, Wildes RP, Jepson AD. P³IV: Probabilistic procedure planning from instructional videos with
   327   328   329   330   331   332   333   334   335   336   337