Page 107 - 《软件学报》2026年第4期
P. 107
软件学报 ISSN 1000-9825, CODEN RUXUEW E-mail: jos@iscas.ac.cn
2026,37(4):1548−1559 [doi: 10.13328/j.cnki.jos.007526] [CSTR: 32375.14.jos.007526] http://www.jos.org.cn
©中国科学院软件研究所版权所有. Tel: +86-10-62562563
*
数值型标签噪声的渐进式区间校正方法
姜高霞 1 , 雷 凡 1 , 张 佳 1 , 王文剑 1,2
1
(山西大学 计算机与信息技术学院, 山西 太原 030006)
2
(数据智能与认知计算山西省重点实验室 (山西大学), 山西 太原 030006)
通信作者: 王文剑, E-mail: wjwang@sxu.edu.cn
摘 要: 在回归任务中, 数值型标签噪声会扭曲数据的真实分布, 削弱模型的泛化能力. 数据过滤是目前常用的一
类方法, 在一定程度上能减少噪声影响, 但易引发过度过滤问题, 导致有效样本流失和数据分布偏移. 提出一种回
归噪声标签的渐进式区间校正 (progressive interval correction, PIC) 算法, 旨在解决数据过滤导致的样本流失问题,
并有效降低标签噪声水平. 首先基于真实标签的后验分布给出标签校正的有效性条件, 以确保降低标签噪声水平;
然后对满足有效性条件的标签进行最大后验校正; 最后通过逐步缩小可信区间范围的方式渐进地校正和优化标签.
在基准数据集与真实数据集上的实验结果表明, PIC 算法能够显著降低数据的噪声水平, 有效提升模型性能.
关键词: 标签噪声; 回归; 标签校正; 渐进式区间校正; 噪声估计
中图法分类号: TP18
中文引用格式: 姜高霞, 雷凡, 张佳, 王文剑. 数值型标签噪声的渐进式区间校正方法. 软件学报, 2026, 37(4): 1548–1559. http://
www.jos.org.cn/1000-9825/7526.htm
英文引用格式: Jiang GX, Lei F, Zhang J, Wang WJ. Progressive Interval Correction Method for Numerical Label Noise. Ruan Jian
Xue Bao/Journal of Software, 2026, 37(4): 1548–1559 (in Chinese). http://www.jos.org.cn/1000-9825/7526.htm
Progressive Interval Correction Method for Numerical Label Noise
1 1 1 1,2
JIANG Gao-Xia , LEI Fan , ZHANG Jia , WANG Wen-Jian
1
(School of Computer and Information Technology, Shanxi University, Taiyuan 030006, China)
2
(Key Laboratory of Data Intelligence and Cognitive Computing of Shanxi Province (Shanxi University), Taiyuan 030006, China)
Abstract: In regression tasks, numerical label noise can distort the true distribution of data and weaken the generalization ability of
models. Data filtering is a commonly used approach that can reduce the impact of noise to some extent. However, it is prone to the issue
of over-filtering, leading to the loss of effective samples and the shift of data distribution. This study presents a progressive interval
correction (PIC) algorithm for regression label noise. The aim is to tackle the problem of sample loss caused by data filtering and
effectively reduce the label noise level. First, based on the posterior distribution of the true labels, the validity conditions for label
correction are established to ensure a reduction in the label noise level. Then, the labels that meet the validity conditions are corrected
using the maximum a posteriori method. Finally, the labels are progressively corrected and optimized by gradually narrowing the range of
the credible interval. Experimental results on both benchmark and real-world datasets demonstrate that the PIC algorithm can significantly
reduce the noise level of data and effectively enhance the performance of models.
Key words: label noise; regression; label correction; progressive interval correction (PIC); noise estimation
1 引 言
高质量的标签数据对于模型的准确训练和良好的泛化能力至关重要. 然而, 从现实世界中获取的数据往往包
* 基金项目: 国家自然科学基金 (62476157, 62276161, U21A20513, 61906113); 山西省基础研究计划 (202303021221055)
本文由“高可信自适应机器学习”专题特约编辑王熙照教授、张敏灵教授、曹付元教授推荐.
收稿时间: 2025-05-12; 修改时间: 2025-06-30, 2025-08-15; 采用时间: 2025-08-20; jos 在线出版时间: 2025-09-02
CNKI 网络首发时间: 2026-01-27

