基于有限数据的高效神经网络黑盒后门检测
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:


Efficient Black-box Backdoor Detection with Limited Data for Neural Networks
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    尽管深度神经网络(deep neural network, DNN)已广泛应用于各个领域, 但其在敌对环境中表现出显著脆弱性. 近年来, 出现一种特殊的攻击形式, 称为后门攻击或木马攻击. 攻击者故意在DNN中植入后门, 带有后门的模型对正常输入做出正确的预测, 但对带有触发器的输入产生错误的预测. 为应对后门攻击的威胁, 研究人员提出各种检测方法, 然而, 这些方法依赖于严格的假设, 如白盒访问目标模型或已知触发器模式, 并且需要大量检测数据和检测时间, 这限制了在实际场景中的应用. 此外, 现有方法在后门攻击的目标类识别方面研究不足, 未能有效揭示目标类在攻击中的关键作用及其对实际任务的影响. 提出一种高效的黑盒后门检测方法 EBLD (efficient black-box backdoor detection), 在仅有少量干净数据的情况下, 结合迁移学习, 利用这些数据训练一小组良性模型和后门模型; 然后通过模型变异技术快速生成大量模型; 最后, 基于这些模型, 训练多个分类器以联合判断目标模型是否为后门模型. 在没有检测数据的情况下, 提出“特征适配”策略, 利用在其他任务上生成的良性模型和后门模型, 训练能够检测新目标模型的二元分类器. 该策略克服了对检测样本的依赖, 并且能够充分利用已有良性模型与后门模型资源, 提升检测效率. 此外, 设计基于迁移学习的二元分类器微调方法, 通过加载已有二元分类器的权重, 训练新的二元分类器, 进一步减少后门检测时间. 最后, 基于训练二元分类器时获得的最优查询集, 构建最优查询集库, 通过分析后门模型对最优查询集的输出类别分布, 快速判断后门攻击的目标类. 目标类识别不仅有助于更精确的定位攻击范围, 还能为后门防御与修复提供明确方向. 实验结果表明, 所提方法在检测精度方面显著优于现有方法, 在检测效率方面也达到较高水平.

    Abstract:

    Although deep neural networks (DNNs) have been widely used in various fields, they exhibit significant vulnerability in adversarial environments. In recent years, a specialized form of attack has emerged, called backdoor attacks or Trojan attacks. Attackers intentionally implant backdoors in DNNs, causing backdoored models to make correct predictions on normal inputs but produce incorrect predictions on inputs containing triggers. To address the threat of backdoor attacks, researchers have proposed various detection methods. However, these methods rely on strict assumptions, such as white-box access to the target model or knowledge of trigger patterns, and require a large amount of detection data and detection time, which limits their applicability in real-world scenarios. Moreover, existing methods have insufficient research on identifying target classes in backdoor attacks and fail to effectively reveal the critical role of target classes in attacks and their impact on real-world tasks. This study proposes an efficient black-box backdoor detection method (EBLD). With only a small amount of clean data, transfer learning is leveraged to train a small set of benign models and backdoored models. Subsequently, a large number of models are rapidly generated through model mutation techniques. Finally, based on these models, multiple classifiers are trained to jointly determine whether the target model is a backdoored model. In the absence of detection data, this study proposes a “feature adaptation” strategy, which leverages benign models and backdoored models generated on other tasks to train binary classifiers capable of detecting new target models. This strategy overcomes the reliance on detection samples and fully utilizes existing benign and backdoored model resources, thus improving detection efficiency. In addition, a binary classifier fine-tuning method based on transfer learning is designed. By loading the weights of existing binary classifiers to train new binary classifiers, backdoor detection time can be further reduced. Finally, an optimal query set library is constructed based on the optimal query sets obtained during binary classifier training. By analyzing the output class distribution of backdoored models on the optimal query sets, the target class of the backdoor attack is quickly identified. Target class identification not only helps to more precisely locate the attack scope but also provides clear guidance for backdoor defense and repair. Experimental results demonstrate that the proposed method significantly outperforms existing methods in detection accuracy and achieves a high level of detection efficiency.

    参考文献
    相似文献
    引证文献
引用本文

李尚松,裴文希,罗润洲,李政辉,刘晓杉,孔维强.基于有限数据的高效神经网络黑盒后门检测.软件学报,,():1-21

复制
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-06-20
  • 最后修改日期:2026-02-26
  • 录用日期:
  • 在线发布日期: 2026-07-08
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62562563 传真:010-62562533 Email:jos@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号