QRCE: 基于查询路由的鲁棒基数估计方法
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

TP311

基金项目:

国家自然科学基金(62462051); 宁夏自然科学基金(2025AAC020045)


QRCE: Robust Cardinality Estimation Method Based on Query Routing
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    基数估计是数据库管理系统(database management system, DBMS)查询优化器的核心组件, 其准确度直接影响执行计划质量. 现有学习式基数估计方法虽然在特定场景下优于传统方法, 但面对高度异质查询负载(如表数量、连接形态及谓词差异巨大)时, 单一模型难以保持稳定的高精度, 同时, 现有多模型方案缺乏精细的查询级选择机制. 因此, 提出基于查询路由的基数估计方法——QRCE, 核心思想包括: 1)构建异质、可扩展的Seq2Seq模型空间, 集成多种具有不同结构归纳偏置的序列建模架构, 形成候选模型集合; 2)构建门控混合查询路由器GMQR, 实现细粒度的、基于查询语义的模型路由, 并体现出面向单条查询的自适应决策特征; 3)提出一种基于Q-error的(1+ε)-近似标签策略, 将模型选择转化为可监督学习任务, 使系统能够根据每条查询的语义特征自适应地选择最适合的模型. 在STATS、JOB-light和TPC-H这3个基准数据集上的实验结果表明, QRCE的整体与尾部误差表现显著优于多数基线方法, 尤其在结构复杂的STATS上, QRCE将Q99误差降低了约53%, 有效规避了长尾灾难性偏差; 在JOB-light上, 其Q99与最优单模型相近, 而在Q50/Q90/Q95上保持优势; 在TPC-H上Q99进一步降低约16%. QRCE在长尾查询和复杂连接场景下展现出更强的鲁棒性与适应性.

    Abstract:

    Cardinality estimation is a core component of the query optimizer in database management systems (DBMS), and its accuracy directly affects the quality of execution plans. Although existing learning-based cardinality estimation methods outperform traditional methods in certain scenarios, they struggle to maintain consistently high accuracy on highly heterogeneous query workloads, where queries may differ significantly in the number of tables, join patterns, and predicates. Moreover, existing multi-model approaches lack a fine-grained mechanism for selecting models at the query level. Therefore, this study proposes a query routing-based cardinality estimation method, QRCE, whose core ideas include 1) constructing a heterogeneous and scalable Seq2Seq model space, integrating multiple sequence modeling architectures with different structural inductive biases to form a candidate model set; 2) building a gated mixture query router, GMQR, to achieve fine-grained model routing based on query semantics and enable adaptive decision-making for individual queries; 3) proposing a Q-error-based (1+ε)-approximate label strategy, transforming model selection into a supervised learning task, enabling the system to adaptively select the most suitable model based on the semantic features of each query. Experimental results on three benchmark datasets, STATS, JOB-light, and TPC-H, show that QRCE achieves significantly better overall and tail error performance than most baseline methods. Particularly on the structurally complex STATS dataset, QRCE reduces the Q99 error by approximately 53%, effectively avoiding long-tail catastrophic bias. On JOB-light, its Q99 is close to that of the best single model, while maintaining advantages at Q50/Q90/Q95. On TPC-H, the Q99 error is further reduced by about 16%. QRCE demonstrates stronger robustness and adaptability in long-tail queries and complex join scenarios.

    参考文献
    相似文献
    引证文献
引用本文

余晓婷,高锦涛. QRCE: 基于查询路由的鲁棒基数估计方法.软件学报,,():1-18

复制
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-01-22
  • 最后修改日期:2026-04-02
  • 录用日期:
  • 在线发布日期: 2026-08-12
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62562563 传真:010-62562533 Email:jos@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号