Page 333 - 《软件学报》2026年第4期
P. 333

1774                                                       软件学报  2026  年第  37  卷第  4  期


                     weak  supervision.  In:  Proc.  of  the  2022  IEEE/CVF  Conf.  on  Computer  Vision  and  Pattern  Recognition.  New  Orleans:  IEEE,  2022.
                     2928–2938. [doi: 10.1109/CVPR52688.2022.00295]
                 [22]   Wang AL, Lin KY, Du JR, Meng JK, Zheng WS. Event-guided procedure planning from instructional videos with text supervision. In:
                     Proc. of the 2023 IEEE/CVF Int’l Conf. on Computer Vision. Paris: IEEE, 2023. 13519–13529. [doi: 10.1109/ICCV51070.2023.01248]
                 [23]   Wang HL, Wu YL, Guo S, Wang LM. PDPP: Projected diffusion for procedure planning in instructional videos. In: Proc. of the 2023
                     IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023. 14836–14845. [doi: 10.1109/CVPR52729.2023.
                     01425]
                 [24]   Li ZH, Geng WJ, Li MH, Chen L, Tang YS, Lu JW, Zhou J. Skip-plan: Procedure planning in instructional videos via condensed action
                     space  learning.  In:  Proc.  of  the  2023  IEEE/CVF  Int’l  Conf.  on  Computer  Vision.  Paris:  IEEE,  2023.  10263–10272.  [doi:  10.1109/
                     ICCV51070.2023.00945]
                 [25]   Zare A, Niu YL, Ayyubi H, Chang SF. RAP: Retrieval-augmented planner for adaptive procedure planning in instructional videos. In:
                     Proc. of the 18th European Conf. on Computer Vision. Milan: Springer, 2024. 410–426. [doi: 10.1007/978-3-031-72980-5_24]
                 [26]   Nagasinghe  KRY,  Zhou  HL,  Gunawardhana  M,  Min  MR,  Harari  D,  Khan  MH.  Why  not  use  your  textbook?  Knowledge-enhanced
                     procedure planning of instructional videos. In: Proc. of the 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Seattle:
                     IEEE, 2024. 18816–18826. [doi: 10.1109/CVPR52733.2024.01780]
                 [27]   Ghallab M, Nau D, Traverso P. Automated Planning: Theory and Practice. San Francisco: Morgan Kaufmann Publishers Inc., 2004.
                 [28]   Viterbi A. Error bounds for convolutional codes and an asymptotically optimum decoding algorithm. IEEE Trans. on Information Theory,
                     1967, 13(2): 260–269. [doi: 10.1109/TIT.1967.1054010]
                 [29]   Fang  F,  Liu  Y,  Koksal  A,  Xu  QL,  Lim  JH.  Masked  diffusion  with  task-awareness  for  procedure  planning  in  instructional  videos.
                     arXiv:2309.07409, 2023.
                 [30]   Shi L, Bürkner PC, Bulling A. ActionDiffusion: An action-aware diffusion model for procedure planning in instructional videos. In: Proc.
                     of the 2025 IEEE/CVF Winter Conf. on Applications of Computer Vision. Tucson: IEEE, 2025. 8816–8825. [doi: 10.1109/WACV61041.
                     2025.00854]
                 [31]   Sun JK, Huang DA, Lu B, Liu YH, Zhou BL, Garg A. PlaTe: Visually-grounded planning with Transformers in procedural tasks. IEEE
                     Robotics and Automation Letters, 2022, 7(2): 4924–4930. [doi: 10.1109/LRA.2022.3150855]
                 [32]   Das P, Xu CL, Doell RF, Corso JJ. A thousand frames in just a few words: Lingual description of videos through latent topics and sparse
                     object stitching. In: Proc. of the 2013 IEEE Conf. on Computer Vision and Pattern Recognition. Portland: IEEE, 2013. 2634–2641. [doi:
                     10.1109/CVPR.2013.340]
                 [33]   Stein S, McKenna SJ. Combining embedded accelerometers with computer vision for recognizing food preparation activities. In: Proc. of
                     the  2013  ACM  Int’l  Joint  Conf.  on  Pervasive  and  Ubiquitous  Computing.  Zurich:  ACM,  2013.  729–738.  [doi:  10.1145/2493432.
                     2493482]
                 [34]   Kuehne H, Arslan A, Serre T. The language of actions: Recovering the syntax and semantics of goal-directed human activities. In: Proc.
                     of the 2014 IEEE Conf. on Computer Vision and Pattern Recognition. Columbus: IEEE, 2014. 780–787. [doi: 10.1109/CVPR.2014.105]
                 [35]   Zhou LW, Xu CL, Corso J. Towards automatic learning of procedures from Web instructional videos. In: Proc. of the 32nd AAAI Conf.
                     on Artificial Intelligence. New Orleans: AAAI Press, 2018. 7590–7598. [doi: 10.1609/AAAI.V32I1.12342]
                 [36]   Rohrbach M, Amin S, Andriluka M, Schiele B. A database for fine grained activity detection of cooking activities. In: Proc. of the 2012
                     IEEE Conf. on Computer Vision and Pattern Recognition. Providence: IEEE, 2012. 1194–1201. [doi: 10.1109/CVPR.2012.6247801]
                 [37]   Kuehne H, Richard A, Gall J. Weakly supervised learning of actions from transcripts. Computer Vision and Image Understanding, 2017,
                     163: 78–89. [doi: 10.1016/J.CVIU.2017.06.004]
                 [38]   Richard A, Kuehne H, Gall J. Action sets: Weakly supervised action segmentation without ordering constraints. In: Proc. of the 2018
                     IEEE Conf. on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018. 5987–5996. [doi: 10.1109/CVPR.2018.00627]
                 [39]   Lu  ZJ,  Elhamifar  E.  Set-supervised  action  learning  in  procedural  task  videos  via  pairwise  order  consistency.  In:  Proc.  of  the  2022
                     IEEE/CVF Conf. on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022. 19871–19881. [doi: 10.1109/CVPR52688.
                     2022.01928]
                 [40]   Elhamifar E, Huynh D. Self-supervised multi-task procedure learning from instructional videos. In: Proc. of the 16th European Conf. on
                     Computer Vision. Glasgow: Springer, 2020. 557–573. [doi: 10.1007/978-3-030-58520-4_33]
                 [41]   Cao M, Yang TY, Weng JW, Zhang C, Wang J, Zou YX. LocVTP: Video-text pre-training for temporal localization. In: Proc. of the 17th
                     European Conf. on Computer Vision. Tel Aviv: Springer, 2022. 38–56. [doi: 10.1007/978-3-031-19809-0_3]
                 [42]   Huang  DA,  Buch  S,  Dery  L,  Garg  A,  Li  FF,  Niebles  JC.  Finding  “it”:  Weakly-supervised  reference-aware  visual  grounding  in
                     instructional  videos.  In:  Proc.  of  the  2018  IEEE  Conf.  on  Computer  Vision  and  Pattern  Recognition.  Salt  Lake  City:  IEEE,  2018.
   328   329   330   331   332   333   334   335   336   337   338