模型互联网中基于置信度的Token级并联协作推理
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

TP18

基金项目:

国家自然科学基金国际合作重点项目(W2411053); 国家自然科学基金联合基金重点项目(U23B2027); 广东省基础与应用基础研究重大项目(2026B0303000007)


Confidence-aware Token-level Parallel Collaborative Inference in AI-model Network
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    在模型互联网中, 单个基模型受限于自身推理能力, 其推理精度往往存在瓶颈. 虽然基模型可以通过与协作模型的Token级并联协作提升精度, 但现有方法多采用全程Token级并联协作策略, 导致较高的Token开销和推理时延. 针对上述挑战, 提出一种基于置信度的协作推理方法CobCo. 该方法以基模型置信度为核心调控信号, 动态调节其推理行为. 当基模型置信度较低时, 引入协作推理以降低推理错误风险, 而在置信度较高时采用独立推理以减少冗余协作开销. 为支撑上述推理调节机制, 设计一种面向Token级并联协作的置信度评估算法, 通过将协作结果作为外部参照刻画基模型在逐步推理过程中的决策优劣, 并基于Beta分布对其置信度进行动态量化. 实验结果表明, 在引入CobCo后, 基模型的推理准确率相较于其独立推理提升了2.4%–27.6%. 与现有全程Token级并联协作方法相比, CobCo在推理准确率损失较小的前提下显著降低了Token开销与推理时延. 特别地, 当基模型与协作模型性能相近时, CobCo能在保持推理准确率不下降的情况下, 将Token开销最高降低约20.8%.

    Abstract:

    The inference accuracy of a single base model is often constrained by its reasoning capacity within an AI-model network. Although Token-level parallel collaboration with collaborative models can effectively improve inference accuracy, existing approaches generally rely on full-process Token-level parallel collaboration, resulting in substantial Token overhead and inference latency. To address these challenges, this study proposes a confidence-aware collaborative inference method, named CobCo. The proposed method leverages the base model’s confidence as a core control signal to regulate its inference behavior dynamically. When the base model has low confidence, collaborative inference is activated to reduce the risk of errors. Conversely, when confidence is sufficiently high, the base model switches to independent inference to reduce redundant collaborative interactions. To support this adaptive inference mechanism, a confidence evaluation algorithm tailored for Token-level parallel collaboration is designed. By treating collaborative outcomes as external references, the proposed method characterizes the decision quality of the base model during step-by-step inference and dynamically quantifies its confidence using a Beta distribution. Experimental results show that integrating CobCo improves the base model’s inference accuracy by 2.4% to 27.6% compared with independent inference. Compared with existing full-process Token-level parallel collaboration methods, CobCo significantly reduces Token overhead and inference latency while incurring only a small loss in inference accuracy. Notably, when the base model and collaborative models exhibit comparable performance, CobCo achieves up to a 20.8% reduction in Token overhead without any degradation in inference accuracy.

    参考文献
    相似文献
    引证文献
引用本文

王建辉,李哲涛,刘忠仁,肖勇.模型互联网中基于置信度的Token级并联协作推理.软件学报,,():1-18

复制
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-01-26
  • 最后修改日期:2026-03-17
  • 录用日期:
  • 在线发布日期: 2026-08-26
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62562563 传真:010-62562533 Email:jos@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号