引用本文:鄢杰斌,祝文涛,刘学林,陈俊杰,钱峰,方玉明.基于最大差异化竞争的通用视觉难样本挖掘.软件学报,2026,37(8):3386-3404
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览 540次   下载 324 本文二维码信息
码上扫一扫!
分享到: 微信 更多
基于最大差异化竞争的通用视觉难样本挖掘
鄢杰斌1,2, 祝文涛1,2, 刘学林1,2, 陈俊杰1,2, 钱峰1,2, 方玉明1,2
1.江西财经大学 计算机与人工智能学院, 江西 南昌 330013;2.多媒体智能处理江西省重点实验室 (江西财经大学), 江西 南昌 330013
摘要:
近年来, 深度学习发展迅速, 在计算机视觉研究中取得了巨大的成功. 在发展过程中, 模型的测试和改进方向是研究者们关注的核心. 然而, 视觉模型比较范式是在封闭数据集上训练(验证)和测试, 然后通过测试结果和真实标签的偏差来获得难样本, 用于反馈模型的问题和改进方向. 这种方式存在的问题包括: 1) 数据集中少量的数据无法真实反映模型的问题; 2)模型预训练等一些操作可能导致数据泄露, 因此展现的性能可能有偏差. 提出基于最大差异化竞争的通用视觉难样本挖掘算法, 自动挖掘真实的难样本, 用于指出模型的问题. 所提算法遵循“通过模型博弈来比较模型”的思想, 联合视觉任务内和多视觉任务间预测结果的“不相似性”优化挖掘潜在的难样本, 旨在以可控的、高效的方式为计算机视觉领域提供新的测试基准. 实验证明, 所构建的测试基准GHS-CV相较于单视觉任务的难样本挖掘(语义分割难样本集SS-C, 显著目标检测难样本集SOD-C)更能暴露出模型的缺陷. 其中, 相较于DeepLabv3+模型在SS-C数据集上的性能, DeepLabv3+在GHS-CV数据集上的mIoU 下降了约 20%; 相较于VST模型在SOD-C 数据集上的性能, VST在GHS-CV数据集上的Fβ下降了约 36%.
关键词:  深度学习  计算机视觉  难样本挖掘  最大差异化竞争  模型评价
DOI:10.13328/j.cnki.jos.007551
分类号:TP391
基金项目:国家自然科学基金(62461028, U24A20220, 62562034, 62402201); 江西省自然科学基金(20243BCE51139, 20232BAB202001, 20252BAC240197)
General Visual Hard Sample Mining Based on Maximum Discrepancy Competition
YAN Jie-Bin1,2, ZHU Wen-Tao1,2, LIU Xue-Lin1,2, CHEN Jun-Jie1,2, QIAN Feng1,2, FANG Yu-Ming1,2
1.School of Computing and Artificial Intelligence, Jiangxi University of Finance and Economics, Nanchang 330013, China;2.Jiangxi Provincial Key Laboratory of Multimedia Intelligent Processing (Jiangxi University of Finance and Economics), Nanchang 330013, China
Abstract:
In recent years, deep learning has developed rapidly and achieved significant success in computer vision, with model evaluation and improvement remaining central concerns for researchers. However, the commonly used model comparison paradigm relies on training (or validation) and testing on closed datasets, and then identifies hard samples based on discrepancies between predictions and ground-truth labels, which provide feedback on model weaknesses and directions for improvement. This paradigm suffers from two major limitations: 1) the limited size and coverage of datasets often fail to faithfully reflect the true weaknesses of models; 2) procedures such as pretraining may introduce data leakage, resulting in potential biases in the demonstrated performance. To address these issues, this study proposes a general visual hard sample mining algorithm based on maximum discrepancy competition, which automatically mines real hard samples to reveal models’ deficiencies. The proposed algorithm follows the principle of “comparing models through competition” and optimizes the discovery of potential hard samples by jointly exploiting the intra-task and cross-task prediction dissimilarities, aiming to provide new test benchmarks for the field of computer vision in a controllable and efficient manner. Experimental results demonstrate that the constructed benchmark named GHS-CV exposes models’ weaknesses more effectively than single-task hard sample benchmarks (i.e., the semantic segmentation hard sample set SS-C and the salient object detection hard sample set SOD-C). Specifically, compared to DeepLabv3+ on SS-C, the mIoU drops by about 20% on GHS-CV, while compared to VST on SOD-C, the Fβ decreases by about 36%.
Key words:  deep learning  computer vision  hard sample mining  maximum discrepancy competition  model evaluation

引用本文:
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览次   下载  
分享到: 微信 更多
摘要:
关键词:  
DOI:
分类号:
基金项目:
Abstract:
Key words: