Page 334 - 《软件学报》2026年第4期
P. 334

吴益露 等: 基于分类检索的操作规划方法                                                            1775


                     5948–5957. [doi: 10.1109/CVPR.2018.00623]
                 [43]   Miech  A,  Alayrac  JB,  Smaira  L,  Laptev  I,  Sivic  J,  Zisserman  A.  End-to-end  learning  of  visual  representations  from  uncurated
                     instructional  videos.  In:  Proc.  of  the  2020  IEEE/CVF  Conf.  on  Computer  Vision  and  Pattern  Recognition.  Seattle:  IEEE,  2020.
                     9876–9886. [doi: 10.1109/CVPR42600.2020.00990]
                 [44]   Zhou HL, Martín-Martín R, Kapadia M, Savarese S, Niebles JC. Procedure-aware pretraining for instructional video understanding. In:
                     Proc. of the 2023 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023. 10727–10738. [doi: 10.1109/
                     CVPR52729.2023.01033]
                 [45]   Zhong YW, Yu LC, Bai Y, Li SW, Yan XT, Li Y. Learning procedure-aware video representation from instructional videos and their
                     narrations. In: Proc. of the 2023 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023. 14825–14835.
                     [doi: 10.1109/CVPR52729.2023.01424]
                 [46]   Mavroudi  E,  Afouras  T,  Torresani  L.  Learning  to  ground  instructional  articles  in  videos  through  narrations.  In:  Proc.  of  the  2023
                     IEEE/CVF Int’l Conf. on Computer Vision. Paris: IEEE, 2023. 15155–15167. [doi: 10.1109/ICCV51070.2023.01395]
                 [47]   Bi J, Luo JB, Xu CL. Procedure planning in instructional videos via contextual modeling and model-based policy learning. In: Proc. of
                     the 2021 IEEE/CVF Int’l Conf. on Computer Vision. Montreal: IEEE, 2021. 15591–15600. [doi: 10.1109/ICCV48922.2021.01532]
                 [48]   Bellman R. A Markovian decision process. Indiana University Mathematics Journal, 1957, 6(4): 679–684. [doi: 10.1512/iumj.1957.6.
                     56038]
                 [49]   Niu  YL,  Guo  WL,  Chen  L,  Lin  XD,  Chang  SF.  SCHEMA:  State  CHangEs  MAtter  for  procedure  planning  in  instructional  videos.
                     arXiv:2403.01599, 2024.
                 [50]   Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I. Attention is all you need. In: Proc. of the
                     31st Int’l Conf. on Neural Information Processing Systems. Long Beach: Curran Associates Inc., 2017. 6000–6010.
                 [51]   Farha YA, Richard A, Gall J. When will you do what?—Anticipating temporal occurrences of activities. In: Proc. of the 2018 IEEE/CVF
                     Conf. on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018. 5343–5352. [doi: 10.1109/CVPR.2018.00560]
                 [52]   Furnari A, Farinella G. What would you expect? Anticipating egocentric actions with rolling-unrolling LSTMs and modality attention. In:
                     Proc. of the 2019 IEEE/CVF Int’l Conf. on Computer Vision. Seoul: IEEE, 2019. 6251–6260. [doi: 10.1109/ICCV.2019.00635]
                 [53]   Liang C, Wang WG, Zhou TF, Yang Y. Visual Abductive Reasoning. In: Proc. of the 2022 IEEE/CVF Conf. on Computer Vision and
                     Pattern Recognition. New Orleans: IEEE, 2022. 15544–15554. [doi: 10.1109/CVPR52688.2022.01512]
                 [54]   Tan C, Yeo CK, Tan C, Fernando B. Inferring past human actions in homes with abductive reasoning. arXiv:2210.13984, 2022.
                 [55]   Chi HG, Lee K, Agarwal N, Xu Y, Ramani K, Choi C. AdamsFormer for spatial action localization in the future. In: Proc. of the 2023
                     IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023. 17885–17895. [doi: 10.1109/CVPR52729.2023.
                     01715]
                 [56]   Sener F, Singhania D, Yao A. Temporal aggregate representations for long-range video understanding. In: Proc. of the 16th European
                     Conf. on Computer Vision. Glasgow: Springer, 2020. 154–171. [doi: 10.1007/978-3-030-58517-4_10]
                 [57]   Girdhar R, Grauman K. Anticipative video Transformer. In: Proc. of the 2021 IEEE/CVF Int’l Conf. on Computer Vision. Montreal:
                     IEEE, 2021. 13485–13495. [doi: 10.1109/ICCV48922.2021.01325]
                 [58]   Stergiou  A,  Damen  D.  The  wisdom  of  crowds:  Temporal  progressive  attention  for  early  action  prediction.  In:  Proc.  of  the  2023
                     IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023. 14709–14719. [doi: 10.1109/CVPR52729.2023.
                     01413]
                 [59]   Zhong ZY, Schneider D, Voit M, Stiefelhagen R, Beyerer J. Anticipative feature fusion Transformer for multi-modal action anticipation.
                     In: Proc. of the 2023 IEEE/CVF Winter Conf. on Applications of Computer Vision. Waikoloa: IEEE, 2023. 6057–6066. [doi: 10.1109/
                     WACV56688.2023.00601]
                 [60]   Damen D, Doughty H, Farinella GM, Fidler S, Furnari A, Kazakos E, Moltisanti D, Munro J, Perrett T, Price W, Wray M. Scaling
                     egocentric vision: The EPIC-KITCHENS dataset. In: Proc. of the 15th European Conf. on Computer Vision. Munich: Springer, 2018.
                     753–771. [doi: 10.1007/978-3-030-01225-0_44]
                 [61]   Ko D, Choi J, Ko J, Noh S, On KW, Kim ES, Kim HJ. Video-text representation learning via differentiable weak temporal alignment. In:
                     Proc. of the 2022 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022. 5006–5015. [doi: 10.1109/
                     CVPR52688.2022.00496]
                 [62]   Xu H, Ghosh G, Huang PY, Okhonko D, Aghajanyan A, Metze F, Zettlemoyer L, Feichtenhofer C. VideoCLIP: Contrastive pre-training
                     for zero-shot video-text understanding. In: Proc. of the 2021 Conf. on Empirical Methods in Natural Language Processing. Punta Cana:
                     ACL, 2021. 6787–6800.
                 [63]   Su WJ, Zhu XZ, Cao Y, Li B, Lu LW, Wei FR, Dai JF. VL-BERT: Pre-training of generic visual-linguistic representations. In: Proc. of
   329   330   331   332   333   334   335   336   337   338   339