基于隐私保护分组生成的高效联邦学习
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

TP311

基金项目:

浙江省尖兵研发攻关计划(2024C01021); 浙江省科技创新领军人才计划(2023R5214)


Efficient Federated Learning with Privacy-preserving Grouping and Generation
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    联邦学习是目前安全多方机器学习领域最为先进的技术, 其因为解决了数据不离开本地的多节点联合模型训练问题而得到了广泛的关注. 然而, 现实中客户端数据的非独立同分布(Non-IID)会显著降低全局模型的性能. 现有的主流方法一类聚焦于在训练时校正模型的参数偏差, 但其在数据分布越极端的情况下效果越差. 另一类方法旨在本地训练前校正数据分布的偏差, 但其会有隐私泄露的风险. 提出一种框架, 遵循严格的联邦学习隐私保护要求, 利用差分隐私、安全多方计算等技术, 在客户端数据分布保持黑盒的前提下, 选择出对全局模型性能提升最有益的若干标签, 针对这些标签生成高质量合成数据并分发, 使补齐这些数据后的客户端本地数据分布趋向于IID, 从而大大提高全局模型的表现. 具体来说, 首先设计一个隐私安全的算法, 对标签在客户端上的分布进行捕获, 并依据该标签分布特征对所有客户端进行组别的划分. 随后挑选出最值得生成的标签, 在标签对应的组内协调客户端合作训练高质量全局生成模型存放于服务器. 最后在任务模型的训练阶段, 基于标签分布特征利用这些生成模型有选择地合成样本并分发, 来使本地数据分布更加均匀, 进而缩小本地模型差距, 在聚合后得到高质量的全局模型. 实验表明, 所提方法能够有效提高全局模型在测试集上的准确率, 且能减少全局模型收敛前所需要的联邦学习通信轮次, 其有效性超越了各基线方法, 在不同的数据集上得到了验证.

    Abstract:

    Federated learning is currently the most advanced technology in secure multi-party machine learning and has attracted significant attention for its ability to enable collaborative model training across multiple nodes without requiring data to leave local devices. However, in real-world scenarios, client data are often non-independent and identically distributed (Non-IID), which significantly degrades the performance of the global model. Current mainstream methods primarily focus on correcting model parameter biases during training, but their effectiveness diminishes as data distribution becomes more extreme. Another category of methods aims to correct data distribution biases before local training, but this approach poses privacy risks. This study proposes a framework that adheres to strict federated learning privacy protection requirements and utilizes techniques such as differential privacy and secure multi-party computation. Assuming that client data distributions remain black boxes, the most beneficial labels for improving global model performance are selected. High-quality synthetic data are generated for these labels and distributed to clients, thus making the local data distribution closer to IID and significantly enhancing the performance of the global model. Specifically, a privacy-preserving algorithm is first designed to capture the distribution of labels across clients, and clients are grouped according to these label distribution characteristics. Then, the most worthwhile labels for generation are selected, and within the corresponding groups, clients are coordinated to collaboratively train high-quality global generative models, which are stored on the server. During the training phase of the task model, based on label distribution characteristics, samples are selectively synthesized using these generative models and distributed to clients, resulting in a more uniform local data distribution. This process reduces local model disparities and yields a high-quality global model after aggregation. Experimental results demonstrate that the proposed method can effectively improve the accuracy of the global model on the test set and reduce the number of federated learning communication rounds required before convergence. Its effectiveness surpasses that of baseline methods and has been validated on different datasets.

    参考文献
    相似文献
    引证文献
引用本文

周众泽,张俊,寿黎但,陈珂,陈刚.基于隐私保护分组生成的高效联邦学习.软件学报,,():1-23

复制
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2024-05-10
  • 最后修改日期:2025-03-13
  • 录用日期:
  • 在线发布日期: 2026-05-20
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62562563 传真:010-62562533 Email:jos@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号