Page 427 - 《软件学报》2026年第5期
P. 427
2306 软件学报 2026 年第 37 卷第 5 期
[13] Yin ZZ, Li BH, Wang M, He RL, Wu WL, Wang HF. Partial multimodal hashing based on fine-grained feature fusion. Ruan Jian Xue
Bao/Journal of Software, 2024, 35(3): 1074–1089 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/7076.htm [doi: 10.
13328/j.cnki.jos.007076]
[14] He J, Chen JN, Liu S, Kortylewski A, Yang C, Bai YT, Wang CH. TransFG: A Transformer architecture for fine-grained recognition. In:
Proc. of the 36th AAAI Conf. on Artificial Intelligence. AAAI Press, 2022. 852–860. [doi: 10.1609/aaai.v36i1.19967]
[15] Sun HB, He XT, Peng YX. SIM-Trans: Structure information modeling Transformer for fine-grained visual categorization. In: Proc. of
the 30th ACM Int’l Conf. on Multimedia. Lisbon: ACM, 2022. 5853–5861. [doi: 10.1145/3503161.3548308]
[16] Xu Q, Wang JH, Jiang B, Luo B. Fine-grained visual classification via internal ensemble learning Transformer. IEEE Trans. on
Multimedia, 2023, 25: 9015–9028. [doi: 10.1109/TMM.2023.3244340]
[17] Wah C, Branson S, Welinder P, Perona P, Belongie S. The Caltech-UCSD birds-200-2011 dataset. 2011. https://gwern.net/doc/ai/dataset/
2011-wah.pdf
[18] Khosla A, Jayadevaprakash N, Yao B, Li FF. Novel dataset for fine-grained image categorization: Stanford Dogs. In: Proc. of the 1st
Workshop on Fine-grained Visual Categorization. Colorado Springs: CVPR, 2011.
[19] van Horn G, Branson S, Farrell R, Haber S, Barry J, Ipeirotis P, Perona P, Belongie S. Building a bird recognition APP and large scale
dataset with citizen scientists: The fine print in fine-grained dataset collection. In: Proc. of the 2015 IEEE Conf. on Computer Vision and
Pattern Recognition. Boston: IEEE, 2015. 595–604. [doi: 10.1109/CVPR.2015.7298658]
[20] van Horn G, Aodha OM, Song Y, Cui Y, Sun C, Shepard A, Adam H, Perona P, Belongie S. The iNaturalist species classification and
detection dataset. In: Proc. of the 2018 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Salt Lake City, 2018. 8769–8778.
[doi: 10.1109/CVPR.2018.00914]
[21] Gavves E, Fernando B, Snoek CGM, Smeulders AWM, Tuytelaars T. Fine-grained categorization by alignments. In: Proc. of the 2013
IEEE Int’l Conf. on Computer Vision. Sydney, 2013. 1713–1720. [doi: 10.1109/ICCV.2013.215]
[22] Branson S, van Horn G, Wah C, Perona P, Belongie S. The ignorant led by the blind: A hybrid human-machine vision system for fine-
grained categorization. Int’l Journal of Computer Vision, 2014, 108(1): 3–29. [doi: 10.1007/s11263-014-0698-4]
[23] Zhang N, Donahue J, Girshick R, Darrell T. Part-based R-CNNs for fine-grained category detection. In: Proc. of the 13th European Conf.
on Computer Vision. Zurich: Springer, 2014. 834–849. [doi: 10.1007/978-3-319-10590-1_54]
[24] Xiao TJ, Xu YC, Yang KY, Zhang JX, Peng YX, Zhang Z. The application of two-level attention models in deep convolutional neural
network for fine-grained image classification. In: Proc. of the 2015 IEEE Conf. on Computer Vision and Pattern Recognition. Boston:
IEEE, 2015. 842–850. [doi: 10.1109/CVPR.2015.7298685]
[25] Fu JL, Zheng HL, Mei T. Look closer to see better: Recurrent attention convolutional neural network for fine-grained image recognition.
In: Proc. of the 2017 IEEE Conf. on Computer Vision and Pattern Recognition. Honolulu: IEEE, 2017. 4476–4484. [doi: 10.1109/CVPR.
2017.476]
[26] Ding Y, Zhou YZ, Zhu Y, Ye QX, Jiao JB. Selective sparse sampling for fine-grained image recognition. In: Proc. of the 2019 IEEE/CVF
Int’l Conf. on Computer Vision. Seoul: IEEE, 2019. 6598–6607. [doi: 10.1109/ICCV.2019.00670]
[27] Ding YF, Ma ZY, Wen SG, Xie JY, Chang DL, Si ZW, Wu M, Ling HB. AP-CNN: Weakly supervised attention pyramid convolutional
neural network for fine-grained visual classification. IEEE Trans. on Image Processing, 2021, 30: 2826–2836. [doi: 10.1109/TIP.2021.
3055617]
[28] Yu CJ, Zhao XY, Zheng Q, Zhang P, You XG. Hierarchical bilinear pooling for fine-grained visual recognition. In: Proc. of the 15th
European Conf on Computer Vision. Munich: Springer, 2018. 595–610. [doi: 10.1007/978-3-030-01270-0_35]
[29] Zhuang PQ, Wang YL, Qiao Y. Learning attentive pairwise interaction for fine-grained classification. In: Proc. of the 34th AAAI Conf.
on Artificial Intelligence. New York: AAAI Press, 2020. 13130–13137. [doi: 10.1609/aaai.v34i07.7016]
[30] Zhang Y, Cao J, Zhang L, Liu XC, Wang ZY, Ling F, Chen WQ. A free lunch from ViT: Adaptive attention multi-scale fusion
Transformer for fine-grained visual recognition. In: Proc. of the 2022 IEEE Int’l Conf. on Acoustics, Speech and Signal Processing.
Singapore, Singapore: IEEE, 2022. 3234–3238. [doi: 10.1109/ICASSP43922.2022.9747591]
[31] He KM, Chen XL, Xie SN, Li YH, Dollár P, Girshick R. Masked autoencoders are scalable vision learners. In: Proc. of the 2022
IEEE/CVF Conf. on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022. 15979–15988. [doi: 10.1109/CVPR52688.
2022.01553]
[32] Xie ZD, Zhang Z, Cao Y, Lin YT, Bao JM, Yao ZL, Dai Q, Hu H. SimMIM: A simple framework for masked image modeling. In: Proc.
of the 2022 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022. 9653–9663. [doi: 10.1109/
CVPR52688.2022.00943]
[33] Tang H, Yuan CC, Li ZC, Tang JH. Learning attention-guided pyramidal features for few-shot fine-grained recognition. Pattern
Recognition, 2022, 130: 108792. [doi: 10.1016/j.patcog.2022.108792]

