Page 57 - 《软件学报》2026年第5期
P. 57

软件学报 ISSN 1000-9825, CODEN RUXUEW                                        E-mail: jos@iscas.ac.cn
                 2026,37(5):1936−1949 [doi: 10.13328/j.cnki.jos.007543] [CSTR: 32375.14.jos.007543]  http://www.jos.org.cn
                 ©中国科学院软件研究所版权所有.                                                          Tel: +86-10-62562563



                                                                      *
                 基于    CLIP    引导标签优化的弱监督图像哈希

                 李泽超  1 ,    金    露  1 ,    王浩骅  1 ,    唐金辉  2


                 1
                  (南京理工大学 计算机科学与工程学院, 江苏 南京 210094)
                 2
                  (南京林业大学, 江苏 南京 210037)
                 通信作者: 金露, E-mail: lu.jin@njust.edu.cn

                 摘 要: 在大规模图像检索任务中, 图像哈希技术通常依赖大量人工标注数据来训练深度哈希模型, 但高昂的人工
                 标注成本限制了其实际应用. 为缓解对人工标注的依赖, 现有研究尝试利用网络用户提供的文本作为弱监督信息,
                 引导模型从图像中挖掘和文本关联的语义信息. 然而, 用户标签中普遍存在噪声, 限制了这些方法的性能. 多模态
                 预训练基础模型      (如  CLIP) 具备较强的图像-文本对齐能力. 受此启发, 利用             CLIP  来优化用户标签, 并提出一种
                 CLIP  引导标签优化的弱监督哈希方法          (CLIP-guided tag refinement hashing, CTRH). 该方法包含  3  个主要内容: 标
                 签置换模块、标签赋权模块和标签平衡损失函数. 标签置换模块通过微调                        CLIP  挖掘图像关联的潜在标签. 标签赋
                 权模块利用优化后的文本和图像进行跨模态全局语义交互, 学习判别性的联合表示. 针对用户标签的分布不平衡
                 问题, 设计了一种标签平衡损失, 通过动态加权增强模型对困难样本的表征学习. 在                         MirFlickr 和  NUS-WIDE  两个
                 通用数据集上与最先进的方法对比验证了所提方法的有效性.
                 关键词: 图像检索; 弱监督哈希; 预训练多模态基础模型; 标签优化
                 中图法分类号: TP391


                 中文引用格式: 李泽超, 金露, 王浩骅, 唐金辉. 基于CLIP引导标签优化的弱监督图像哈希. 软件学报, 2026, 37(5): 1936–1949. http://
                 www.jos.org.cn/1000-9825/7543.htm
                 英文引用格式: Li ZC, Jin L, Wang HH, Tang JH. Weakly Supervised Image Hashing via CLIP-guided Tag Refinement. Ruan Jian Xue
                 Bao/Journal of Software, 2026, 37(5): 1936–1949 (in Chinese). http://www.jos.org.cn/1000-9825/7543.htm
                 Weakly Supervised Image Hashing via CLIP-guided Tag Refinement

                         1
                                              1
                                1
                 LI Ze-Chao , JIN Lu , WANG Hao-Hua , TANG Jin-Hui 2
                 1
                 (School of Computer Science and Engineering, Nanjing University of Science and Technology, Nanjing 210094, China)
                 2
                 (Nanjing Forestry University, Nanjing 210037, China)
                 Abstract:  In  large-scale  image  retrieval  tasks,  image  hashing  typically  relies  on  a  large  amount  of  manually  annotated  data  to  train  deep
                 hashing  models.  However,  the  high  cost  of  manual  annotation  limits  its  practical  application.  To  alleviate  this  dependency,  existing  studies
                 attempt  to  use  texts  provided  by  Web  users  as  weak  supervision  to  guide  the  model  in  mining  semantic  information  associated  with  the
                 texts  from  images.  Nevertheless,  the  inherent  noise  in  user  tags  often  limits  model  performance.  Multimodal  pre-trained  models  such  as
                 CLIP  exhibit  strong  image-text  alignment  capabilities.  Inspired  by  this,  this  study  utilizes  CLIP  to  optimize  user  tags  and  proposes  a
                 weakly  supervised  hashing  method  called  CLIP-guided  tag  refinement  hashing  (CTRH).  The  proposed  method  consists  of  three  key
                 components:  a  tag  replacement  module,  a  tag  weighting  module,  and  a  tag-balanced  loss  function.  The  tag  replacement  module  fine-tunes
                 CLIP  to  mine  potential  image-relevant  tags.  The  tag  weighting  module  performs  cross-modal  global  semantic  interaction  between  the
                 optimized  text  and  images  to  learn  discriminative  joint  representations.  To  address  the  imbalance  of  user  tags,  a  tag-balanced  loss  is


                 *    基金项目: 国家自然科学基金  (62425603, 62372233); 江苏省基础研究计划攀登项目  (BK20240011)
                  本文由“多媒体智能理解与生成”专题特约编辑孙立峰教授、闵巍庆副研究员、马占宇教授、蒋树强研究员、彭宇新教授、田丰研究
                  员、黄庆明教授推荐.
                  收稿时间: 2025-05-26; 修改时间: 2025-07-11; 采用时间: 2025-09-05; jos 在线出版时间: 2025-09-23
                  CNKI 网络首发时间: 2026-01-08
   52   53   54   55   56   57   58   59   60   61   62