FedProbe: 代理模型引导的可解释联邦学习投毒攻击防御框架
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

北京市自然科学基金(M21039)


FedProbe: Proxy-model-guided Explainable Defense Framework for Poisoning Attacks in Federated Learning
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    联邦学习允许多个参与方在不共享本地私有数据的前提下协同训练一个共享的深度学习模型. 然而, 该范式在实际应用中极易受到日益复杂的投毒攻击威胁. 现有的防御方法在检测效率、对多样化攻击的泛化能力以及在非独立同分布 (non-IID) 数据环境下的性能稳定性方面仍存在局限. 为应对上述挑战, 提出一种名为 FedProbe 的、由代理模型引导的可解释联邦学习投毒攻击防御框架. FedProbe采用两阶段机制: 首先, 在代理模型引导阶段, 利用服务器端的可信样本集计算各本地模型间的KL散度以度量其行为相似性, 并确定最优聚类数量. 在剔除离群更新后, 各簇被聚合成代理模型, 此举旨在有效划分由不同数据分布训练的模型并提升后续检测效率. 其次, 在可解释分析阶段, FedProbe利用SHAP可解释技术对代理模型进行深度分析, 计算其关键特征归因, 并最终通过跨类别可疑度分数来精准识别并剔除潜在的恶意模型. 实验结果表明, FedProbe在多种基准数据集上展现出卓越的性能与鲁棒性. 在良性环境中, 其收敛时间仅为标准聚合算法的1.11–1.21倍, 性能开销极小. 在安全性方面, 即使在恶意用户占比高达40%的极端场景下, FedProbe在抵御无目标攻击和后门攻击时仍能保持优越的收敛性, 并将攻击成功率控制在25%以下; 而在面对自适应攻击时, 其在防御有效性与任务准确率方面均具有显著优势.

    Abstract:

    Federated Learning enables collaborative model training among multiple clients without exchanging their local private data. However, this paradigm is vulnerable to increasingly sophisticated poisoning attacks. Existing defense mechanisms often exhibit shortcomings in detection efficiency, generalization to diverse attacks, and performance stability under non-independent and identically distributed (non-IID) data. To address these challenges, this study proposes FedProbe, an interpretable federated poisoning attack defense framework guided by proxy models. FedProbe adopts a two-stage mechanism. In the first stage, under proxy model guidance, the framework computes the KL divergence among local models on a trusted server-side dataset to measure their behavioral similarity and determine the optimal number of clusters. After removing outlier updates, each cluster is aggregated into a proxy model, aiming to effectively partition models trained on different data distributions and improve subsequent detection efficiency. In the second stage, an interpretable analysis is conducted. FedProbe leverages the SHAP technique to perform an in-depth analysis of the proxy models, computing key feature attributions and ultimately using cross-category suspicion scores to identify and remove potentially malicious models accurately. Experimental results demonstrate that FedProbe exhibits superior performance and robustness across various benchmark datasets. In benign settings, the convergence time is only 1.11 to 1.21 times that of standard aggregation algorithms, indicating minimal overhead. In terms of security, even in extreme scenarios with up to 40% malicious clients, FedProbe maintains superior convergence when defending against untargeted attacks and backdoor attacks, while keeping the attack success rate below 25%. Moreover, when facing adaptive attacks, it shows significant advantages in both defense effectiveness and task accuracy.

    参考文献
    相似文献
    引证文献
引用本文

曹益皓,张建标,赵亚茹,韩宇飞,张承昱,叶涛. FedProbe: 代理模型引导的可解释联邦学习投毒攻击防御框架.软件学报,,():1-20

复制
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-08-06
  • 最后修改日期:2025-12-30
  • 录用日期:
  • 在线发布日期: 2026-06-03
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62562563 传真:010-62562533 Email:jos@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号