交互式蒸馏驱动的开源项目领域分类方法
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

TP311

基金项目:

国家自然科学基金重点项目(62032016); 湖北省自然科学基金青年项目(2023AFB374); 南京大学计算机软件新技术全国重点实验室开放课题(KFKT2025B48)


Interactive-distillation-driven Method for Open-source Project Domain Classification
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    随着开源社区的发展, 开源软件项目的种类也在逐步增多. 但开源社区中已有的数据, 例如标签(tag)、描述等不能直接推断出项目的细分领域, 造成了开发者和用户需要花费非必要的时间去人为判断开源软件项目的类别、开源研究者研究的项目样本中可能无法有效区分细粒度类别的项目, 进而影响研究结果的针对性. 针对以上问题, 构建交互式蒸馏方法, 可以对开源项目进行自动化分类工作, 其中类别包括代码开发工具或插件、网页应用等12个类别. 首先, 标注了30310个开源软件项目. 然后, 对已有的大模型方法以及机器学习方法进行项目细粒度领域分类任务测试. 最后, 基于基准方法存在的问题, 利用大模型蒸馏关键词特征, 再通过引入大小模型协同方法优化模型分类能力. 在标注的数据集上的验证结果显示: 基于交互式蒸馏的开源项目细粒度分类方法在十折交叉后取得了92.1%的平均精确率、91.8%的平均准确率和91.8%的平均F1值, 相较于基准方法提升了26.3%–32.8%的精确率, 24.5%–39.9%的准确率, 26.1%–40.7%的F1值.

    Abstract:

    With the continuous development of open-source communities, the variety of open-source software projects has also been steadily increasing. However, existing data in these communities, such as tags and descriptions, cannot be used to directly infer the fine-grained domains of projects. As a result, developers and users need to spend unnecessary time manually identifying the categories of open-source software projects, and open-source researchers may be unable to effectively distinguish fine-grained categories in their project samples, thus affecting the specificity of research results. To address these issues, this study proposes an interactive distillation method that can automatically classify open-source projects into 12 categories, including code development tools or plugins, Web applications, and other categories. First, 30310 open-source software projects are annotated in this study. Then, existing large model methods and machine learning methods are tested on the fine-grained domain classification task. Finally, based on the problems identified in the baseline methods, this study distills keyword features from large models and further optimizes the model’s classification capability by introducing a collaboration method between large and small models. Validation results on the annotated dataset show that the proposed method achieves an average Precision of 92.1%, an average Accuracy of 91.8%, and an average F1-measure of 91.8% after ten-fold cross-validation. Compared with the baseline methods, the proposed method improves Precision by 26.3%–32.8%, Accuracy by 24.5%–39.9%, and F1-measure by 26.1%–40.7%.

    参考文献
    相似文献
    引证文献
引用本文

康成希,程璨,张能,游兰,李兵,王伟.交互式蒸馏驱动的开源项目领域分类方法.软件学报,,():1-24

复制
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-09-26
  • 最后修改日期:2026-02-02
  • 录用日期:
  • 在线发布日期: 2026-08-19
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62562563 传真:010-62562533 Email:jos@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号