引用本文:黄小龙,杨婧如,柳熠,马郓,景翔,黄罡.ReproLink: 面向可复现性的科研数据管理系统.软件学报,2025,36(12):5801-5820
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览 1067次   下载 1809 本文二维码信息
码上扫一扫!
分享到: 微信 更多
ReproLink: 面向可复现性的科研数据管理系统
黄小龙1,2, 杨婧如2,3, 柳熠2,3, 马郓2,4, 景翔2,5, 黄罡1,2
1.北京大学 计算机学院, 北京 100871;2.数据空间技术与系统全国重点实验室, 北京 100091;3.北京大数据先进技术研究院, 北京 100195;4.北京大学 人工智能研究院, 北京 100871;5.北京大学 软件与微电子学院, 北京 102600
摘要:
科研成果的可复现性是科学研究可靠性的基本保证, 更是科学技术进步的基石. 然而, 当前学术界面临着严峻的可复现性危机, 大量在顶级期刊和会议上公开发表的科研成果无法复现. 在数据科学领域, 成果的可复现性面临着科研数据多源异构、计算流程复杂、计算环境复杂等挑战. 针对这些问题, 提出面向可复现性的科研数据管理系统ReproLink. ReproLink提出对科研数据的统一建模, 将科研数据抽象为包含标识、属性集、数据实体三要素的科研数据对象; 通过对于复现流程的细粒度建模, ReproLink建立一种对多步骤复杂复现流程的精确描述方法. 通过代码和运行环境的一体化建模, ReproLink消除不同环境中代码执行行为的不确定性给成果复现带来的影响. 对ReproLink的性能测试和实例分析表明, ReproLink在百万级的数据规模下具有较好的性能表现, 在论文复现、复现相关数据的溯源等现实场景中具有实用价值. ReproLink系统技术架构已集成到国内唯一专门面向科研院所的一体化综合管理与服务平台-科南软件, 支持国内数百家科研机构的成果复现需求.
关键词:  科研数据管理  可复现性  数字对象架构  数据语用  数据共享
DOI:10.13328/j.cnki.jos.007372
分类号:TP311
基金项目:北京市科技新星计划(Z211100002121159); 数据空间技术与系统全国重点实验室资助项目
ReproLink: Reproducibility-oriented Research Data Management System
HUANG Xiao-Long1,2, YANG Jing-Ru2,3, LIU Yi2,3, MA Yun2,4, JING Xiang2,5, HUANG Gang1,2
1.School of Computer Science, Peking University, Beijing 100871, China;2.National Key Laboratory of Data Space Technology and System, Beijing 100091, China;3.Advanced Institute of Big Data Technology, Beijing 100195, China;4.Institute for Artificial Intelligence, Peking University, Beijing 100871, China;5.School of Software and Microelectronics, Peking University, Beijing 102600, China
Abstract:
The reproducibility of scientific research results is a fundamental guarantee for the reliability of scientific research and the cornerstone of scientific and technological advancement. However, the research community is currently facing a serious reproducibility crisis, with many research results published in top journals and conferences being irreproducible. In the field of data science, the reproducibility of research results faces challenges such as heterogeneous research data from multiple sources, complex computational processes, and intricate computational environments. To address these issues, this study proposes ReproLink, a reproducibility-oriented research data management system. ReproLink constructs a unified model of research data, abstracting it into research data objects that consist of three elements: identifier, attribute set, and data entity. Through fine-grained modeling of the reproduction process, ReproLink establishes a precise method for describing multi-step, complex reproduction processes. By integrating code and operating environment modeling, ReproLink eliminates the uncertainties caused by different environments affecting code execution. Performance tests and case studies show that ReproLink performs well with data scales up to one million records, demonstrating practical value in real-world scenarios such as paper reproduction and data provenance tracking. The technical architecture of ReproLink has been integrated into Conow Software, the only integrated comprehensive management and service platform in China specifically designed for scientific research institutes, supporting the reproducibility needs of hundreds of such institutes across the country.
Key words:  scientific data management  reproducibility  digital object architecture (DOA)  data provenance  data sharing

引用本文:
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览次   下载  
分享到: 微信 更多
摘要:
关键词:  
DOI:
分类号:
基金项目:
Abstract:
Key words: