Page 57 - 《软件学报》2026年第5期
P. 57
软件学报 ISSN 1000-9825, CODEN RUXUEW E-mail: jos@iscas.ac.cn
2026,37(5):1936−1949 [doi: 10.13328/j.cnki.jos.007543] [CSTR: 32375.14.jos.007543] http://www.jos.org.cn
©中国科学院软件研究所版权所有. Tel: +86-10-62562563
*
基于 CLIP 引导标签优化的弱监督图像哈希
李泽超 1 , 金 露 1 , 王浩骅 1 , 唐金辉 2
1
(南京理工大学 计算机科学与工程学院, 江苏 南京 210094)
2
(南京林业大学, 江苏 南京 210037)
通信作者: 金露, E-mail: lu.jin@njust.edu.cn
摘 要: 在大规模图像检索任务中, 图像哈希技术通常依赖大量人工标注数据来训练深度哈希模型, 但高昂的人工
标注成本限制了其实际应用. 为缓解对人工标注的依赖, 现有研究尝试利用网络用户提供的文本作为弱监督信息,
引导模型从图像中挖掘和文本关联的语义信息. 然而, 用户标签中普遍存在噪声, 限制了这些方法的性能. 多模态
预训练基础模型 (如 CLIP) 具备较强的图像-文本对齐能力. 受此启发, 利用 CLIP 来优化用户标签, 并提出一种
CLIP 引导标签优化的弱监督哈希方法 (CLIP-guided tag refinement hashing, CTRH). 该方法包含 3 个主要内容: 标
签置换模块、标签赋权模块和标签平衡损失函数. 标签置换模块通过微调 CLIP 挖掘图像关联的潜在标签. 标签赋
权模块利用优化后的文本和图像进行跨模态全局语义交互, 学习判别性的联合表示. 针对用户标签的分布不平衡
问题, 设计了一种标签平衡损失, 通过动态加权增强模型对困难样本的表征学习. 在 MirFlickr 和 NUS-WIDE 两个
通用数据集上与最先进的方法对比验证了所提方法的有效性.
关键词: 图像检索; 弱监督哈希; 预训练多模态基础模型; 标签优化
中图法分类号: TP391
中文引用格式: 李泽超, 金露, 王浩骅, 唐金辉. 基于CLIP引导标签优化的弱监督图像哈希. 软件学报, 2026, 37(5): 1936–1949. http://
www.jos.org.cn/1000-9825/7543.htm
英文引用格式: Li ZC, Jin L, Wang HH, Tang JH. Weakly Supervised Image Hashing via CLIP-guided Tag Refinement. Ruan Jian Xue
Bao/Journal of Software, 2026, 37(5): 1936–1949 (in Chinese). http://www.jos.org.cn/1000-9825/7543.htm
Weakly Supervised Image Hashing via CLIP-guided Tag Refinement
1
1
1
LI Ze-Chao , JIN Lu , WANG Hao-Hua , TANG Jin-Hui 2
1
(School of Computer Science and Engineering, Nanjing University of Science and Technology, Nanjing 210094, China)
2
(Nanjing Forestry University, Nanjing 210037, China)
Abstract: In large-scale image retrieval tasks, image hashing typically relies on a large amount of manually annotated data to train deep
hashing models. However, the high cost of manual annotation limits its practical application. To alleviate this dependency, existing studies
attempt to use texts provided by Web users as weak supervision to guide the model in mining semantic information associated with the
texts from images. Nevertheless, the inherent noise in user tags often limits model performance. Multimodal pre-trained models such as
CLIP exhibit strong image-text alignment capabilities. Inspired by this, this study utilizes CLIP to optimize user tags and proposes a
weakly supervised hashing method called CLIP-guided tag refinement hashing (CTRH). The proposed method consists of three key
components: a tag replacement module, a tag weighting module, and a tag-balanced loss function. The tag replacement module fine-tunes
CLIP to mine potential image-relevant tags. The tag weighting module performs cross-modal global semantic interaction between the
optimized text and images to learn discriminative joint representations. To address the imbalance of user tags, a tag-balanced loss is
* 基金项目: 国家自然科学基金 (62425603, 62372233); 江苏省基础研究计划攀登项目 (BK20240011)
本文由“多媒体智能理解与生成”专题特约编辑孙立峰教授、闵巍庆副研究员、马占宇教授、蒋树强研究员、彭宇新教授、田丰研究
员、黄庆明教授推荐.
收稿时间: 2025-05-26; 修改时间: 2025-07-11; 采用时间: 2025-09-05; jos 在线出版时间: 2025-09-23
CNKI 网络首发时间: 2026-01-08

