Page 69 - 《软件学报》2026年第5期
P. 69
1948 软件学报 2026 年第 37 卷第 5 期
哈希码学习联合特征蕴含的语义知识. 在 MirFlickr 和 NUS-WIDE 上的实验结果验证了本文所提方法的有效性.
References
[1] Huang XY, Sun B, Yang ZY, Zhu YY, Tian Q. Locality-sensitive hashing approach based on semantic space for visual retrieval. Journal
of Image and Graphics, 2021, 26(7): 1568–1582 (in Chinese with English abstract). [doi: 10.11834/jig.200534]
[2] Li ZX, Ling F, Tang ZJ, Ma HF, Shi ZP. Unsupervised cross-media hashing retrieval based on multi-head attention network. Scientia
Sinica Informationis, 2021, 51(12): 2053–2068 (in Chinese with English abstract). [doi: 10.1360/SSI-2020-0264]
[3] Xia RK, Pan Y, Lai HJ, Liu C, Yan SC. Supervised hashing for image retrieval via image representation learning. In: Proc. of the 28th
AAAI Conf. on Artificial Intelligence. Québec City: AAAI, 2014. 2156–2162.
[4] Hussain A, Li HC, Ali D, Ali M, Abbas F, Hussain M. An optimized deep supervised hashing model for fast image retrieval. Image and
Vision Computing, 2023, 133: 104668. [doi: 10.1016/j.imavis.2023.104668]
[5] Yang F, Ding XJ, Liu YF, Ma FM, Cao J. Scalable semantic-enhanced supervised hashing for cross-modal retrieval. Knowledge-based
Systems, 2022, 251: 109176. [doi: 10.1016/j.knosys.2022.109176]
[6] Shi NF, Fu C, Tie M, Zhang WC, Wang XW, Sham CW. Attention-based deep supervised hashing for near duplicate video retrieval.
Neural Computing and Applications, 2024, 36(10): 5217–5230. [doi: 10.1007/s00521-023-09342-x]
[7] Karaman S, Lin XD, Hu XF, Chang SF. Unsupervised rank-preserving hashing for large-scale image retrieval. In: Proc. of the 2019 Int’l
Conf. on Multimedia Retrieval. Ottawa: ACM, 2019. 192–196. [doi: 10.1145/3323873.3325038]
[8] Xiong SY, Pan LL, Ma XQ, Hu QH, Beckman E. Unsupervised deep hashing with multiple similarity preservation for cross-modal image-
text retrieval. Int’l Journal of Machine Learning and Cybernetics, 2024, 15(10): 4423–4434. [doi: 10.1007/s13042-024-02154-y]
[9] Yang Z, Deng XY, Long J. Fast unsupervised consistent and modality-specific hashing for multimedia retrieval. Neural Computing &
Applications, 2023, 35(8): 6207–6223. [doi: 10.1007/s00521-022-08008-4]
[10] Venkateswara H, Eusebio J, Chakraborty S, Panchanathan S. Deep hashing network for unsupervised domain adaptation. In: Proc. of the
2017 IEEE Conf. on Computer Vision and Pattern Recognition. Honolulu: IEEE, 2017. 5385–5394. [doi: 10.1109/CVPR.2017.572]
[11] Gattupalli V, Zhuo YX, Li BX. Weakly supervised deep image hashing through tag embeddings. In: Proc. of the 2019 IEEE/CVF Conf.
on Computer Vision and Pattern Recognition. Long Beach: IEEE, 2019. 10367–10376. [doi: 10.1109/CVPR.2019.01062]
[12] Jin L, Li ZC, Pan YH, Tang JH. Weakly-supervised image hashing through masked visual-semantic graph-based reasoning. In: Proc. of
the 28th ACM Int’l Conf. on Multimedia. Seattle: ACM, 2020. 916–924. [doi: 10.1145/3394171.3414022]
[13] Wang M, Zhou WG, Tian Q, Li HQ. Deep enhanced weakly-supervised hashing with iterative tag refinement. IEEE Trans. on
Multimedia, 2022, 24: 2779–2790. [doi: 10.1109/TMM.2021.3087356]
[14] Du YC, Wang M, Lu ZB, Zhou WG, Li HQ. Weakly supervised hashing with reconstructive cross-modal attention. ACM Trans. on
Multimedia Computing, Communications and Applications, 2023, 19(6): 208. [doi: 10.1145/3589185]
[15] Kim W, Son B, Kim I. ViLT: Vision-and-language Transformer without convolution or region supervision. In: Proc. of the 38th Int’l
Conf. on Machine Learning. Online: PMLR, 2021. 5583–5594.
[16] Chen YC, Li LJ, Yu LC, El Kholy A, Ahmed F, Gan Z, Cheng Y, Liu JJ. UNITER: Universal image-text representation learning. In:
Proc. of the 16th European Conf. on Computer Vision. Glasgow: Springer, 2020. 104–120. [doi: 10.1007/978-3-030-58577-8_7]
[17] Li XJ, Yin X, Li CY, Zhang PC, Hu XW, Zhang L, Wang LJ, Hu HD, Dong L, Wei FR, Choi YJ, Gao JF. OSCAR: Object-semantics
aligned pre-training for vision-language tasks. In: Proc. of the 16th European Conf. on Computer Vision. Glasgow: Springer, 2020.
121–137. [doi: 10.1007/978-3-030-58577-8_8]
[18] Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, Krueger G, Sutskever I. Learning
transferable visual models from natural language supervision. In: Proc. of the 38th Int’l Conf. on Machine Learning. PMLR, 2021.
8748–8763.
[19] Chatfield K, Simonyan K, Vedaldi A, Zisserman A. Return of the devil in the details: Delving deep into convolutional nets. arXiv:
1405.3531, 2014.
[20] Mikolov T, Chen K, Corrado G, Dean J. Efficient estimation of word representations in vector space. arXiv:1301.3781, 2013.
[21] Yin J, Zhang ZD, Gao YH, Yang ZW, Li L, Xiao M, Sun YQ, Yan CG. Survey on vision-language pre-training. Ruan Jian Xue
Bao/Journal of Software, 2023, 34(5): 2000–2023 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/6774.htm [doi: 10.
13328/j.cnki.jos.006774]
[22] Zhang HY, Wang TB, Li MZ, Zhao Z, Pu SL, Wu F. Comprehensive review of visual-language-oriented multimodal pre-training
methods. Journal of Image and Graphics, 2022, 27(9): 2652–2682 (in Chinese with English abstract). [doi: 10.11834/jig.220173]
[23] Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I. Attention is all you need. In: Proc. of the

