基于数据重建的净标签数据投毒攻击方案
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

国家自然科学基金(U25A20429, 62332018, 62332004); 中央高校基本科研业务费专项资金 (ZYGX2025TS010); 四川省自然科学基金(2025ZNSFSC0497); 深圳市科技重大专项(KJZD20231023092900002)


Clean-label Data Poisoning Attack Scheme Based on Data Reconstruction
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    人工智能系统依赖大规模、开放式的训练数据采集, 这为数据投毒攻击提供了可乘之机. 攻击者能够通过渗透数据供应链等方式操控训练数据, 破坏深度学习模型的可用性. 相较于基于标签翻转的脏标签数据投毒攻击, 净标签攻击虽具有更强的隐蔽性, 但在实现时面临更大挑战: 一方面, 需要精确操控样本特征而非直接篡改其标签; 另一方面, 还需确保毒化样本在视觉上与原始样本保持高度一致以规避检测. 虽然现有的净标签攻击方案已取得一定成效, 但往往因假设宽松, 难以在实际场景中有效实施, 且隐蔽性和计算效率仍有提升空间. 为此, 提出一种基于数据重建的深度学习净标签数据投毒攻击方案CPDR, 该方案重构原用于图像分割任务的Attention U-Net模型的损失函数, 将其转换为一个包含重建损失、投毒效果损失和模型更新一致性损失的多任务学习重建模型, 通过数据重建生成隐蔽的毒化样本, 并在黑盒假设下实现有效的投毒攻击. 在集中式学习和联邦学习这2种场景、2个数据集和4个目标模型上的实验表明, CPDR能在几乎不影响深度学习主任务表现的情况下降低模型在目标类别上的分类准确率. 在集中式学习下, CPDR对目标类别准确率的平均降幅高于VagueGAN方案2.33个百分点, 高于Adversarial Poison方案2.42个百分点; 在联邦学习下, 其优势进一步扩大, 分别高于上述对比方案4.44和4.10个百分点. 此外, CPDR生成的毒化样本与原始样本视觉差异微小, 难以辨别和检测, 实现了隐蔽性和攻击效果的良好平衡.

    Abstract:

    Artificial intelligence systems rely on large-scale, open training data collection, which creates opportunities for data poisoning attacks. Attackers can manipulate training data by infiltrating the data supply chain, thus compromising the availability of deep learning models. Compared to dirty-label data poisoning attacks based on label flipping, clean-label attacks are more stealthy but face greater challenges in implementation. On the one hand, precise manipulation of sample features is required rather than direct alteration of labels. On the other hand, it is necessary to ensure that poisoned samples remain highly visually consistent with the original samples to evade detection. Although existing clean-label attacks have achieved some success, they are often difficult to implement in practical scenarios due to loose assumptions, and there remains room for improvement in terms of stealthiness and computational efficiency. To address these issues, this study proposes a clean-label data poisoning attack scheme for deep learning based on data reconstruction, termed CPDR. In this scheme, the loss function of the Attention U-Net model, originally designed for image segmentation, is reconstructed and transformed into a multi-task learning reconstruction model that incorporates reconstruction loss, attack effectiveness loss, and model update consistency loss. Through data reconstruction, stealthy poisoned samples are generated, and effective poisoning attacks are achieved under a black-box assumption. Experiments conducted in both centralized and federated learning scenarios, across two datasets and four target models, demonstrate that CPDR reduces the model’s classification accuracy on target classes with minimal impact on the performance of the main task. In centralized learning, CPDR achieves an additional reduction of 2.33 and 2.42 percentage points in target class accuracy compared to VagueGAN and Adversarial Poison, respectively. In federated learning, this advantage further increases to 4.44 and 4.10 percentage points, respectively. Moreover, the poisoned samples generated by CPDR exhibit minimal visual differences from the original samples, making them difficult to distinguish and detect, thus achieving a good balance between stealthiness and attack effectiveness.

    参考文献
    相似文献
    引证文献
引用本文

张文博,李雄,张文琪,谢勇,陈厅,禹继国,张小松.基于数据重建的净标签数据投毒攻击方案.软件学报,,():1-20

复制
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-07-07
  • 最后修改日期:2025-09-18
  • 录用日期:
  • 在线发布日期: 2026-07-08
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62562563 传真:010-62562533 Email:jos@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号