Page 69 - 《软件学报》2026年第5期
P. 69

1948                                                       软件学报  2026  年第  37  卷第  5  期


                 哈希码学习联合特征蕴含的语义知识. 在             MirFlickr 和  NUS-WIDE  上的实验结果验证了本文所提方法的有效性.

                 References
                  [1]   Huang XY, Sun B, Yang ZY, Zhu YY, Tian Q. Locality-sensitive hashing approach based on semantic space for visual retrieval. Journal
                     of Image and Graphics, 2021, 26(7): 1568–1582 (in Chinese with English abstract). [doi: 10.11834/jig.200534]
                  [2]   Li ZX, Ling F, Tang ZJ, Ma HF, Shi ZP. Unsupervised cross-media hashing retrieval based on multi-head attention network. Scientia
                     Sinica Informationis, 2021, 51(12): 2053–2068 (in Chinese with English abstract). [doi: 10.1360/SSI-2020-0264]
                  [3]   Xia RK, Pan Y, Lai HJ, Liu C, Yan SC. Supervised hashing for image retrieval via image representation learning. In: Proc. of the 28th
                     AAAI Conf. on Artificial Intelligence. Québec City: AAAI, 2014. 2156–2162.
                  [4]   Hussain A, Li HC, Ali D, Ali M, Abbas F, Hussain M. An optimized deep supervised hashing model for fast image retrieval. Image and
                     Vision Computing, 2023, 133: 104668. [doi: 10.1016/j.imavis.2023.104668]
                  [5]   Yang F, Ding XJ, Liu YF, Ma FM, Cao J. Scalable semantic-enhanced supervised hashing for cross-modal retrieval. Knowledge-based
                     Systems, 2022, 251: 109176. [doi: 10.1016/j.knosys.2022.109176]
                  [6]   Shi NF, Fu C, Tie M, Zhang WC, Wang XW, Sham CW. Attention-based deep supervised hashing for near duplicate video retrieval.
                     Neural Computing and Applications, 2024, 36(10): 5217–5230. [doi: 10.1007/s00521-023-09342-x]
                  [7]   Karaman S, Lin XD, Hu XF, Chang SF. Unsupervised rank-preserving hashing for large-scale image retrieval. In: Proc. of the 2019 Int’l
                     Conf. on Multimedia Retrieval. Ottawa: ACM, 2019. 192–196. [doi: 10.1145/3323873.3325038]
                  [8]   Xiong SY, Pan LL, Ma XQ, Hu QH, Beckman E. Unsupervised deep hashing with multiple similarity preservation for cross-modal image-
                     text retrieval. Int’l Journal of Machine Learning and Cybernetics, 2024, 15(10): 4423–4434. [doi: 10.1007/s13042-024-02154-y]
                  [9]   Yang Z, Deng XY, Long J. Fast unsupervised consistent and modality-specific hashing for multimedia retrieval. Neural Computing &
                     Applications, 2023, 35(8): 6207–6223. [doi: 10.1007/s00521-022-08008-4]
                 [10]   Venkateswara H, Eusebio J, Chakraborty S, Panchanathan S. Deep hashing network for unsupervised domain adaptation. In: Proc. of the
                     2017 IEEE Conf. on Computer Vision and Pattern Recognition. Honolulu: IEEE, 2017. 5385–5394. [doi: 10.1109/CVPR.2017.572]
                 [11]   Gattupalli V, Zhuo YX, Li BX. Weakly supervised deep image hashing through tag embeddings. In: Proc. of the 2019 IEEE/CVF Conf.
                     on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2019. 10367–10376. [doi: 10.1109/CVPR.2019.01062]
                 [12]   Jin L, Li ZC, Pan YH, Tang JH. Weakly-supervised image hashing through masked visual-semantic graph-based reasoning. In: Proc. of
                     the 28th ACM Int’l Conf. on Multimedia. Seattle: ACM, 2020. 916–924. [doi: 10.1145/3394171.3414022]
                 [13]   Wang  M,  Zhou  WG,  Tian  Q,  Li  HQ.  Deep  enhanced  weakly-supervised  hashing  with  iterative  tag  refinement.  IEEE  Trans.  on
                     Multimedia, 2022, 24: 2779–2790. [doi: 10.1109/TMM.2021.3087356]
                 [14]   Du YC, Wang M, Lu ZB, Zhou WG, Li HQ. Weakly supervised hashing with reconstructive cross-modal attention. ACM Trans. on
                     Multimedia Computing, Communications and Applications, 2023, 19(6): 208. [doi: 10.1145/3589185]
                 [15]   Kim W, Son B, Kim I. ViLT: Vision-and-language Transformer without convolution or region supervision. In: Proc. of the 38th Int’l
                     Conf. on Machine Learning. Online: PMLR, 2021. 5583–5594.
                 [16]   Chen YC, Li LJ, Yu LC, El Kholy A, Ahmed F, Gan Z, Cheng Y, Liu JJ. UNITER: Universal image-text representation learning. In:
                     Proc. of the 16th European Conf. on Computer Vision. Glasgow: Springer, 2020. 104–120. [doi: 10.1007/978-3-030-58577-8_7]
                 [17]   Li XJ, Yin X, Li CY, Zhang PC, Hu XW, Zhang L, Wang LJ, Hu HD, Dong L, Wei FR, Choi YJ, Gao JF. OSCAR: Object-semantics
                     aligned  pre-training  for  vision-language  tasks.  In:  Proc.  of  the  16th  European  Conf.  on  Computer  Vision.  Glasgow:  Springer,  2020.
                     121–137. [doi: 10.1007/978-3-030-58577-8_8]
                 [18]   Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, Krueger G, Sutskever I. Learning
                     transferable  visual  models  from  natural  language  supervision.  In:  Proc.  of  the  38th  Int’l  Conf.  on  Machine  Learning.  PMLR,  2021.
                     8748–8763.
                 [19]   Chatfield  K,  Simonyan  K,  Vedaldi  A,  Zisserman  A.  Return  of  the  devil  in  the  details:  Delving  deep  into  convolutional  nets.  arXiv:
                     1405.3531, 2014.
                 [20]   Mikolov T, Chen K, Corrado G, Dean J. Efficient estimation of word representations in vector space. arXiv:1301.3781, 2013.
                 [21]   Yin  J,  Zhang  ZD,  Gao  YH,  Yang  ZW,  Li  L,  Xiao  M,  Sun  YQ,  Yan  CG.  Survey  on  vision-language  pre-training.  Ruan  Jian  Xue
                     Bao/Journal of Software, 2023, 34(5): 2000–2023 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/6774.htm [doi: 10.
                     13328/j.cnki.jos.006774]
                 [22]   Zhang  HY,  Wang  TB,  Li  MZ,  Zhao  Z,  Pu  SL,  Wu  F.  Comprehensive  review  of  visual-language-oriented  multimodal  pre-training
                     methods. Journal of Image and Graphics, 2022, 27(9): 2652–2682 (in Chinese with English abstract). [doi: 10.11834/jig.220173]
                 [23]   Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I. Attention is all you need. In: Proc. of the
   64   65   66   67   68   69   70   71   72   73   74