Root Cause Change Identification for Failures in Large-scale Online System
Author:
Affiliation:

Clc Number:

TP311

Fund Project:

  • Article
  • |
  • Figures
  • |
  • Metrics
  • |
  • Reference
  • |
  • Related
  • |
  • Cited by
  • |
  • Materials
  • |
  • Comments
    Abstract:

    In large-scale online service systems, software changes occur frequently and are on the rise due to the need to adapt to rapidly changing user demands and information technologies, such as continuous integration and delivery. Although engineers rigorously test new software versions before deployment, some defects may go unnoticed during testing and end up being deployed to the production environment. This primarily occurs because significant differences exist between the testing and production environments in terms of load, scale, and user characteristics. As a result, these defects can impact the system’s availability and stability. To better understand the impact and behavior of defective changes after deployment to the production environment, this study conducts an empirical analysis using real change failure data from WeChat, a large-scale global instant messaging system. Five key findings related to defective changes are derived from this analysis. Based on these empirical findings and conclusions, this study proposes a lightweight root cause change identification method. This method aims to automatically identify the root cause changes that lead to failure, assisting operations and maintenance engineers in root cause localization and trouble shooting efforts. To validate the effectiveness of the proposed method, a real dataset containing various types of defective changes from WeChat’s production environment is collected, along with a simulated change dataset based on a microservice benchmark system. A systematic evaluation of the proposed method is then conducted. The experimental results show that the proposed method achieves Top-3 root cause change hit rates of 80% and 84% for the WeChat production environment dataset and simulated change data, respectively, significantly outperforming the state-of-the-art defective change detection methods. Moreover, from an engineering practice perspective, the system uses only 2.3 GB of memory and has an average analysis latency of 28.6 s when processing typical-scale failures, thus meeting the requirements of actual production environments.

    Reference
    Related
    Cited by
Get Citation

余广坝,陈鹏飞,唐锡涛,郑子彬.面向大规模在线系统的故障根因变更识别.软件学报,2026,37(2):641-661

Copy
Share
Article Metrics
  • Abstract:
  • PDF:
  • HTML:
  • Cited by:
History
  • Received:June 07,2024
  • Revised:March 25,2025
  • Adopted:
  • Online: December 03,2025
  • Published: February 06,2026
You are the firstVisitors
Copyright: Institute of Software, Chinese Academy of Sciences Beijing ICP No. 05046678-4
Address:4# South Fourth Street, Zhong Guan Cun, Beijing 100190,Postal Code:100190
Phone:010-62562563 Fax:010-62562533 Email:jos@iscas.ac.cn
Technical Support:Beijing Qinyun Technology Development Co., Ltd.

Beijing Public Network Security No. 11040202500063