不可学习样本生成与净化方法综述
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

TP18

基金项目:

国家自然科学基金(62502221, U22B2062, 62172232); 江苏省自然科学基金(BK20250737); 江苏省科技重大专项(BG2024042); 江苏省高等学校自然科学研究基金(24KJB520019); 南京市科技重大专项(202405002); 南京市留学归国人员科技创新计划(R2024LZ04); 南京信息工程大学人才启动经费(2025r069)


Review of Generation and Purification Methods for Unlearnable Examples
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    大规模深度学习模型的迅速迭代对高质量训练数据的需求不断攀升, 而这一需求通常依赖于从互联网收集海量数据, 其中不乏涉及个人图像与文本的私有信息, 因而引发潜在的隐私泄漏与伦理风险. 为缓解此类问题, 研究者提出了不可学习样本这一主动数据隐私保护策略. 该方法通过向训练数据中嵌入难以察觉的微小扰动, 使未授权模型在表面上获得较高训练精度, 却因学习到无意义的“捷径特征”而失去泛化能力. 对不可学习样本的研究进展进行系统梳理, 将现有工作划分为两大方向: 一类聚焦于不可学习样本的生成方法, 另一类则关注其净化与检测机制. 在此基础上, 进一步探讨未来的研究挑战与发展路径, 旨在为隐私数据保护提供理论支撑与技术突破.

    Abstract:

    The rapid iteration of large-scale deep learning models leads to an ever-growing demand for high-quality training data. This demand is typically met by collecting massive amounts of data from the Internet, which often include private information such as personal images and text, thus creating risks related to privacy leakage and ethical concerns. To mitigate such issues, researchers have proposed unlearnable examples as a proactive strategy for data privacy protection. This approach embeds imperceptible perturbations into training data, making unauthorized models appear to achieve high training accuracy while actually losing generalization ability due to learning meaningless “shortcut features.” This study provides a systematic review of recent advances in unlearnable examples, categorizing existing work into two main areas: generation methods and purification and detection mechanisms. Building on this foundation, future research challenges and potential directions are further explored, with the aim of providing theoretical support and facilitating technical breakthroughs in private data protection.

    参考文献
    相似文献
    引证文献
引用本文

孟若涵,卞新玉,张丛阳,朱浩天,付章杰.不可学习样本生成与净化方法综述.软件学报,,():1-30

复制
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-09-26
  • 最后修改日期:2026-04-27
  • 录用日期:
  • 在线发布日期: 2026-08-19
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62562563 传真:010-62562533 Email:jos@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号