Page 428 - 《软件学报》2026年第5期
P. 428

唐昊 等: 基于视觉    Transformer 的双视图融合细粒度图像识别                                         2307


                 [34]   Zha ZC, Tang H, Sun YL, Tang JH. Boosting few-shot fine-grained recognition with background suppression and foreground alignment.
                     IEEE Trans. on Circuits and Systems for Video Technology, 2023, 33(8): 3947–3961. [doi: 10.1109/TCSVT.2023.3236636]
                 [35]   Nair V, Hinton GE. Rectified linear units improve restricted Boltzmann machines. In: Proc. of the 27th Int’l Conf. on Machine Learning.
                     Haifa: Omnipress, 2010. 807–814.
                 [36]   He KM, Zhang XY, Ren SQ, Sun J. Deep residual learning for image recognition. In: Proc. of the 2016 IEEE Conf. on Computer Vision
                     and Pattern Recognition. Las Vegas: IEEE, 2016. 770–778. [doi: 10.1109/CVPR.2016.90]
                 [37]   Zheng HL, Fu JL, Mei T, Luo JB. Learning multi-attention convolutional neural network for fine-grained image recognition. In: Proc. of
                     the 2017 IEEE Int’l Conf. on Computer Vision. Venice: IEEE, 2017. 5209–5217. [doi: 10.1109/ICCV.2017.557]
                 [38]   Yang Z, Luo TG, Wang D, Hu ZQ, Gao J, Wang LW. Learning to navigate for fine-grained classification. In: Proc. of the 15th European
                     Conf. on Computer Vision. Munich: Springer, 2018. 420–435. [doi: 10.1007/978-3-030-01264-9_26]
                 [39]   Luo W, Yang XT, Mo XJ, Lu YH, Davis L, Li J, Yang J, Lim SN. Cross-X learning for fine-grained visual categorization. In: Proc. of the
                     2019 IEEE/CVF Int’l Conf. on Computer Vision. Seoul: IEEE, 2019. 8241–8250. [doi: 10.1109/ICCV.2019.00833]
                 [40]   Chen  Y,  Bai  YL,  Zhang  W,  Mei  T.  Destruction  and  construction  learning  for  fine-grained  image  recognition.  In:  Proc.  of  the  2019
                     IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2019. 5152–5161. [doi: 10.1109/CVPR.2019.00530]
                 [41]   Xu KR, Lai R, Gu L, Li YS. Multiresolution discriminative mixup network for fine-grained visual categorization. IEEE Trans. on Neural
                     Networks and Learning Systems, 2023, 34(7): 3488–3500. [doi: 10.1109/TNNLS.2021.3112768]
                 [42]   Liu  CB,  Xie  HT,  Zha  ZJ,  Ma  LF,  Yu  LY,  Zhang  YD.  Filtration  and  distillation:  Enhancing  region  attention  for  fine-grained  visual
                     categorization. In: Proc. of the 34th AAAI Conf on Artificial Intelligence. New York: AAAI Press, 2020. 11555–11562. [DOI: 10.1609/
                     aaai.v34i07.6822]
                 [43]   Du RY, Xie JY, Ma ZY, Chang DL, Song YZ, Guo J. Progressive learning of category-consistent multi-granularity features for fine-
                     grained visual classification. IEEE Trans. on Pattern Analysis and Machine Intelligence, 2022, 44(12): 9521–9535. [doi: 10.1109/TPAMI.
                     2021.3126668]
                 [44]   Rao YM, Chen GY, Lu JW, Zhou J. Counterfactual attention learning for fine-grained visual categorization and re-identification. In: Proc.
                     of the 2021 IEEE/CVF Int’l Conf. on Computer Vision. Montreal: IEEE, 2021. 1005–1014. [doi: 10.1109/ICCV48922.2021.00106]
                 [45]   Wang Q, Wang JJ, Deng HY, Wu X, Wang YZ, Hao GF. AA-Trans: Core attention aggregating Transformer with information entropy
                     selector for fine-grained visual classification. Pattern Recognition, 2023, 140: 109547. [doi: 10.1016/j.patcog.2023.109547]
                 [46]   Zhang ZC, Chen ZD, Wang YX, Luo X, Xu XS. A vision Transformer for fine-grained classification by reducing noise and enhancing
                     discriminative information. Pattern Recognition, 2024, 145: 109979. [doi: 10.1016/j.patcog.2023.109979]
                 [47]   Jiang X, Tang H, Gao JY, Du XY, He SF, Li ZC. Delving into multimodal prompting for fine-grained visual classification. In: Proc. of
                     the 38th AAAI Conf. on Artificial Intelligence. Vancouver: AAAI Press, 2024. 2570–2578. [doi: 10.1609/aaai.v38i3.28034]
                 [48]   Wang CM, Fu HY, Ma HD. Multi-part token Transformer with dual contrastive learning for fine-grained image classification. In: Proc. of
                     the 31st ACM Int’l Conf. on Multimedia. Ottawa: ACM, 2023. 7648–7656. [doi: 10.1145/3581783.3612303]
                 [49]   Bi Q, Zhou BC, Ji W, Xia GS. Universal fine-grained visual categorization by concept guided learning. IEEE Trans. on Image Processing,
                     2025, 34: 394–409. [doi: 10.1109/TIP.2024.3523802]
                 [50]   Guo P, Farrell R. Aligned to the object, not to the image: A unified pose-aligned representation for fine-grained recognition. In: Proc. of
                     the 2019 IEEE Winter Conf. on Applications of Computer Vision. Waikoloa: IEEE, 2019. 1876–1885. [doi: 10.1109/WACV.2019.00204]
                 [51]   Zhao YF, Yan K, Huang FY, Li J. Graph-based high-order relation discovery for fine-grained recognition. In: Proc. of 2021 IEEE/CVF
                     Conf. on Computer Vision and Pattern Recognition. Nashville: IEEE, 2021. 15074–15083. [doi: 10.1109/CVPR46437.2021.01483]
                 [52]   Cui Y, Song Y, Sun C, Howard A, Belongie S. Large scale fine-grained categorization and domain-specific transfer learning. In: Proc. of
                     the 2018 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018. 4109–4118. [doi: 10.1109/CVPR.
                     2018.00432]
                 [53]   Zhang LB, Huang SL, Liu W, Tao DC. Learning a mixture of granularity-specific experts for fine-grained categorization. In: Proc. of the
                     2019 IEEE/CVF Int’l Conf. on Computer Vision. Seoul: IEEE, 2019. 8330–8339. [doi: 10.1109/ICCV.2019.00842]
                 [54]   Recasens A, Kellnhofer P, Stent S, Matusik W, Torralba A. Learning to zoom: A saliency-based sampling layer for neural networks. In:
                     Proc. of the 15th European Conf. on Computer Vision. Munich: Springer, 2018. 52–67. [doi: 10.1007/978-3-030-01240-3_4]
                 [55]   Huang ZX, Li Y. Interpretable and accurate fine-grained recognition via region grouping. In: Proc. of the 2020 IEEE/CVF Conf. on
                     Computer Vision and Pattern Recognition. Seattle: IEEE, 2020. 8659–8669. [doi: 10.1109/CVPR42600.2020.00869]
                 [56]   Zheng HL, Fu JL, Zha ZJ, Luo JB. Looking for the devil in the details: Learning trilinear attention sampling network for fine-grained
                     image  recognition.  In:  Proc.  of  the  2019  IEEE/CVF  Conf.  on  Computer  Vision  and  Pattern  Recognition.  Long  Beach:  IEEE,  2019.
                     5007–5016. [doi: 10.1109/CVPR.2019.00515]
                 [57]   Wang SJ, Wang ZH, Li HJ, Chang JL, Ouyang WL, Tian Q. Accurate fine-grained object recognition with structure-driven relation graph
   423   424   425   426   427   428   429   430   431   432   433