Page 333 - 《软件学报》2026年第4期
P. 333
1774 软件学报 2026 年第 37 卷第 4 期
weak supervision. In: Proc. of the 2022 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022.
2928–2938. [doi: 10.1109/CVPR52688.2022.00295]
[22] Wang AL, Lin KY, Du JR, Meng JK, Zheng WS. Event-guided procedure planning from instructional videos with text supervision. In:
Proc. of the 2023 IEEE/CVF Int’l Conf. on Computer Vision. Paris: IEEE, 2023. 13519–13529. [doi: 10.1109/ICCV51070.2023.01248]
[23] Wang HL, Wu YL, Guo S, Wang LM. PDPP: Projected diffusion for procedure planning in instructional videos. In: Proc. of the 2023
IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023. 14836–14845. [doi: 10.1109/CVPR52729.2023.
01425]
[24] Li ZH, Geng WJ, Li MH, Chen L, Tang YS, Lu JW, Zhou J. Skip-plan: Procedure planning in instructional videos via condensed action
space learning. In: Proc. of the 2023 IEEE/CVF Int’l Conf. on Computer Vision. Paris: IEEE, 2023. 10263–10272. [doi: 10.1109/
ICCV51070.2023.00945]
[25] Zare A, Niu YL, Ayyubi H, Chang SF. RAP: Retrieval-augmented planner for adaptive procedure planning in instructional videos. In:
Proc. of the 18th European Conf. on Computer Vision. Milan: Springer, 2024. 410–426. [doi: 10.1007/978-3-031-72980-5_24]
[26] Nagasinghe KRY, Zhou HL, Gunawardhana M, Min MR, Harari D, Khan MH. Why not use your textbook? Knowledge-enhanced
procedure planning of instructional videos. In: Proc. of the 2024 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Seattle:
IEEE, 2024. 18816–18826. [doi: 10.1109/CVPR52733.2024.01780]
[27] Ghallab M, Nau D, Traverso P. Automated Planning: Theory and Practice. San Francisco: Morgan Kaufmann Publishers Inc., 2004.
[28] Viterbi A. Error bounds for convolutional codes and an asymptotically optimum decoding algorithm. IEEE Trans. on Information Theory,
1967, 13(2): 260–269. [doi: 10.1109/TIT.1967.1054010]
[29] Fang F, Liu Y, Koksal A, Xu QL, Lim JH. Masked diffusion with task-awareness for procedure planning in instructional videos.
arXiv:2309.07409, 2023.
[30] Shi L, Bürkner PC, Bulling A. ActionDiffusion: An action-aware diffusion model for procedure planning in instructional videos. In: Proc.
of the 2025 IEEE/CVF Winter Conf. on Applications of Computer Vision. Tucson: IEEE, 2025. 8816–8825. [doi: 10.1109/WACV61041.
2025.00854]
[31] Sun JK, Huang DA, Lu B, Liu YH, Zhou BL, Garg A. PlaTe: Visually-grounded planning with Transformers in procedural tasks. IEEE
Robotics and Automation Letters, 2022, 7(2): 4924–4930. [doi: 10.1109/LRA.2022.3150855]
[32] Das P, Xu CL, Doell RF, Corso JJ. A thousand frames in just a few words: Lingual description of videos through latent topics and sparse
object stitching. In: Proc. of the 2013 IEEE Conf. on Computer Vision and Pattern Recognition. Portland: IEEE, 2013. 2634–2641. [doi:
10.1109/CVPR.2013.340]
[33] Stein S, McKenna SJ. Combining embedded accelerometers with computer vision for recognizing food preparation activities. In: Proc. of
the 2013 ACM Int’l Joint Conf. on Pervasive and Ubiquitous Computing. Zurich: ACM, 2013. 729–738. [doi: 10.1145/2493432.
2493482]
[34] Kuehne H, Arslan A, Serre T. The language of actions: Recovering the syntax and semantics of goal-directed human activities. In: Proc.
of the 2014 IEEE Conf. on Computer Vision and Pattern Recognition. Columbus: IEEE, 2014. 780–787. [doi: 10.1109/CVPR.2014.105]
[35] Zhou LW, Xu CL, Corso J. Towards automatic learning of procedures from Web instructional videos. In: Proc. of the 32nd AAAI Conf.
on Artificial Intelligence. New Orleans: AAAI Press, 2018. 7590–7598. [doi: 10.1609/AAAI.V32I1.12342]
[36] Rohrbach M, Amin S, Andriluka M, Schiele B. A database for fine grained activity detection of cooking activities. In: Proc. of the 2012
IEEE Conf. on Computer Vision and Pattern Recognition. Providence: IEEE, 2012. 1194–1201. [doi: 10.1109/CVPR.2012.6247801]
[37] Kuehne H, Richard A, Gall J. Weakly supervised learning of actions from transcripts. Computer Vision and Image Understanding, 2017,
163: 78–89. [doi: 10.1016/J.CVIU.2017.06.004]
[38] Richard A, Kuehne H, Gall J. Action sets: Weakly supervised action segmentation without ordering constraints. In: Proc. of the 2018
IEEE Conf. on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018. 5987–5996. [doi: 10.1109/CVPR.2018.00627]
[39] Lu ZJ, Elhamifar E. Set-supervised action learning in procedural task videos via pairwise order consistency. In: Proc. of the 2022
IEEE/CVF Conf. on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022. 19871–19881. [doi: 10.1109/CVPR52688.
2022.01928]
[40] Elhamifar E, Huynh D. Self-supervised multi-task procedure learning from instructional videos. In: Proc. of the 16th European Conf. on
Computer Vision. Glasgow: Springer, 2020. 557–573. [doi: 10.1007/978-3-030-58520-4_33]
[41] Cao M, Yang TY, Weng JW, Zhang C, Wang J, Zou YX. LocVTP: Video-text pre-training for temporal localization. In: Proc. of the 17th
European Conf. on Computer Vision. Tel Aviv: Springer, 2022. 38–56. [doi: 10.1007/978-3-031-19809-0_3]
[42] Huang DA, Buch S, Dery L, Garg A, Li FF, Niebles JC. Finding “it”: Weakly-supervised reference-aware visual grounding in
instructional videos. In: Proc. of the 2018 IEEE Conf. on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018.

