图数据库节点标识存储引擎性能建模与适配策略
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

TP311

基金项目:

国家自然科学基金(U25B2018, 62322213, 62461146205); 中国博士后科学基金(2025M781495); 国家资助博士后研究人员计划(GZC20251044); 北京市科技计划(Z251100008125032); 清华大学水木学者项目(2025SM061)


Performance Modeling and Adaptation Strategy for Node Identifier Store Engines in Graph Databases
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    随着大数据和人工智能技术的快速发展, 图数据库因其在复杂关系建模与高效查询方面的优势, 逐渐成为社交网络分析、金融风控、知识图谱等领域的核心基础设施. 图数据库的管理对象是节点以及节点之间的关系. 在架构层面, 节点唯一标识符(node identifier, NodeID)作为图数据管理的核心纽带, 承担节点身份表征、关系寻址和图算法执行的关键职能. 当前主流图数据库普遍采用键值存储(key-value store, KVS)引擎实现节点标识到图结构的映射管理. 然而, 现有系统多依赖通用键值存储引擎(如RocksDB)管理此类映射, 却缺乏对负载特性的深度考量, 具体表现为: 1) 节点标识映射负载特征建模缺失; 2) 跨软硬件环境(如CPU/内存、SSD/HDD)的适应性不足. 首先系统分析图数据库中节点标识映射的操作特性与键值存储需求, 进而评估多种主流键值存储引擎(包括RocksDB、LMDB、LevelDB、FasterKV及ForestDB)在异构硬件环境下的性能表现, 系统揭示不同数据负载 (如数据规模、读写比)与硬件配置(如内存容量、线程数、存储介质)对执行效率的影响规律. 基于大规模实验(覆盖5类数据集、3种硬件平台及1300+组对照测试), 提出一种基于决策树模型的适配策略, 整合负载特征(数据规模、读写比)与硬件配置(内存、线程数、磁盘类型), 以指导键值存储引擎的自适应选择. 实验表明, 该模型推荐最优引擎的准确率达92.1%, 次优场景性能差距小于10%.

    Abstract:

    With the rapid development of big data and artificial intelligence technologies, graph databases have gradually become core infrastructure for social network analysis, financial risk control, and knowledge graphs due to their advantages in complex relationship modeling and efficient querying. Graph databases manage two primary objects: nodes and the relationships between them. At the architectural level, the node identifier (NodeID) serves as the critical link for graph data management, undertaking key functions including node identity representation, relationship lookup, and graph algorithm execution. Current mainstream graph databases widely adopt key-value store (KVS) to implement NodeID-to-graph-structure mapping management. However, existing systems largely rely on general-purpose KVSs (e.g., RocksDB) for managing such mappings, yet lack in-depth consideration of workload characteristics. These limitations are reflected in two aspects: (1) the lack of workload modeling for node identifier mapping, and (2) insufficient adaptability to heterogeneous software and hardware environments (e.g., CPU/memory, SSD/HDD). This study first systematically analyzes the operational characteristics and KVS requirements of node identifier mapping in graph databases. It then evaluates the performance of multiple mainstream KVS engines (including RocksDB, LMDB, LevelDB, FasterKV, and ForestDB) in heterogeneous hardware environments, systematically revealing the impact patterns of data workloads (e.g., data scale, read-write ratio) and hardware configurations (e.g., memory capacity, thread count, storage medium) on execution efficiency. Based on large-scale experiments involving five datasets, three hardware platforms, and over 1 300 comparative tests, this study proposes an adaptation strategy based on decision tree model that integrates workload characteristics (data scale and read-write ratio) with hardware configurations (e.g., memory, thread count, disk type) to guide adaptive selection of KVS engines. Experiments show that the model achieves 92.1% accuracy in recommending optimal engines, with suboptimal scenarios exhibiting less than 10% performance gap.

    参考文献
    相似文献
    引证文献
引用本文

陈政,张峰,齐畅,赵郑诣隆,杜小勇.图数据库节点标识存储引擎性能建模与适配策略.软件学报,,():1-23

复制
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-08-18
  • 最后修改日期:2025-11-20
  • 录用日期:
  • 在线发布日期: 2026-06-01
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62562563 传真:010-62562533 Email:jos@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号