Page 334 - 《软件学报》2026年第4期
P. 334
吴益露 等: 基于分类检索的操作规划方法 1775
5948–5957. [doi: 10.1109/CVPR.2018.00623]
[43] Miech A, Alayrac JB, Smaira L, Laptev I, Sivic J, Zisserman A. End-to-end learning of visual representations from uncurated
instructional videos. In: Proc. of the 2020 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Seattle: IEEE, 2020.
9876–9886. [doi: 10.1109/CVPR42600.2020.00990]
[44] Zhou HL, Martín-Martín R, Kapadia M, Savarese S, Niebles JC. Procedure-aware pretraining for instructional video understanding. In:
Proc. of the 2023 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023. 10727–10738. [doi: 10.1109/
CVPR52729.2023.01033]
[45] Zhong YW, Yu LC, Bai Y, Li SW, Yan XT, Li Y. Learning procedure-aware video representation from instructional videos and their
narrations. In: Proc. of the 2023 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023. 14825–14835.
[doi: 10.1109/CVPR52729.2023.01424]
[46] Mavroudi E, Afouras T, Torresani L. Learning to ground instructional articles in videos through narrations. In: Proc. of the 2023
IEEE/CVF Int’l Conf. on Computer Vision. Paris: IEEE, 2023. 15155–15167. [doi: 10.1109/ICCV51070.2023.01395]
[47] Bi J, Luo JB, Xu CL. Procedure planning in instructional videos via contextual modeling and model-based policy learning. In: Proc. of
the 2021 IEEE/CVF Int’l Conf. on Computer Vision. Montreal: IEEE, 2021. 15591–15600. [doi: 10.1109/ICCV48922.2021.01532]
[48] Bellman R. A Markovian decision process. Indiana University Mathematics Journal, 1957, 6(4): 679–684. [doi: 10.1512/iumj.1957.6.
56038]
[49] Niu YL, Guo WL, Chen L, Lin XD, Chang SF. SCHEMA: State CHangEs MAtter for procedure planning in instructional videos.
arXiv:2403.01599, 2024.
[50] Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I. Attention is all you need. In: Proc. of the
31st Int’l Conf. on Neural Information Processing Systems. Long Beach: Curran Associates Inc., 2017. 6000–6010.
[51] Farha YA, Richard A, Gall J. When will you do what?—Anticipating temporal occurrences of activities. In: Proc. of the 2018 IEEE/CVF
Conf. on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018. 5343–5352. [doi: 10.1109/CVPR.2018.00560]
[52] Furnari A, Farinella G. What would you expect? Anticipating egocentric actions with rolling-unrolling LSTMs and modality attention. In:
Proc. of the 2019 IEEE/CVF Int’l Conf. on Computer Vision. Seoul: IEEE, 2019. 6251–6260. [doi: 10.1109/ICCV.2019.00635]
[53] Liang C, Wang WG, Zhou TF, Yang Y. Visual Abductive Reasoning. In: Proc. of the 2022 IEEE/CVF Conf. on Computer Vision and
Pattern Recognition. New Orleans: IEEE, 2022. 15544–15554. [doi: 10.1109/CVPR52688.2022.01512]
[54] Tan C, Yeo CK, Tan C, Fernando B. Inferring past human actions in homes with abductive reasoning. arXiv:2210.13984, 2022.
[55] Chi HG, Lee K, Agarwal N, Xu Y, Ramani K, Choi C. AdamsFormer for spatial action localization in the future. In: Proc. of the 2023
IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023. 17885–17895. [doi: 10.1109/CVPR52729.2023.
01715]
[56] Sener F, Singhania D, Yao A. Temporal aggregate representations for long-range video understanding. In: Proc. of the 16th European
Conf. on Computer Vision. Glasgow: Springer, 2020. 154–171. [doi: 10.1007/978-3-030-58517-4_10]
[57] Girdhar R, Grauman K. Anticipative video Transformer. In: Proc. of the 2021 IEEE/CVF Int’l Conf. on Computer Vision. Montreal:
IEEE, 2021. 13485–13495. [doi: 10.1109/ICCV48922.2021.01325]
[58] Stergiou A, Damen D. The wisdom of crowds: Temporal progressive attention for early action prediction. In: Proc. of the 2023
IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023. 14709–14719. [doi: 10.1109/CVPR52729.2023.
01413]
[59] Zhong ZY, Schneider D, Voit M, Stiefelhagen R, Beyerer J. Anticipative feature fusion Transformer for multi-modal action anticipation.
In: Proc. of the 2023 IEEE/CVF Winter Conf. on Applications of Computer Vision. Waikoloa: IEEE, 2023. 6057–6066. [doi: 10.1109/
WACV56688.2023.00601]
[60] Damen D, Doughty H, Farinella GM, Fidler S, Furnari A, Kazakos E, Moltisanti D, Munro J, Perrett T, Price W, Wray M. Scaling
egocentric vision: The EPIC-KITCHENS dataset. In: Proc. of the 15th European Conf. on Computer Vision. Munich: Springer, 2018.
753–771. [doi: 10.1007/978-3-030-01225-0_44]
[61] Ko D, Choi J, Ko J, Noh S, On KW, Kim ES, Kim HJ. Video-text representation learning via differentiable weak temporal alignment. In:
Proc. of the 2022 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022. 5006–5015. [doi: 10.1109/
CVPR52688.2022.00496]
[62] Xu H, Ghosh G, Huang PY, Okhonko D, Aghajanyan A, Metze F, Zettlemoyer L, Feichtenhofer C. VideoCLIP: Contrastive pre-training
for zero-shot video-text understanding. In: Proc. of the 2021 Conf. on Empirical Methods in Natural Language Processing. Punta Cana:
ACL, 2021. 6787–6800.
[63] Su WJ, Zhu XZ, Cao Y, Li B, Lu LW, Wei FR, Dai JF. VL-BERT: Pre-training of generic visual-linguistic representations. In: Proc. of

