Abstract:As a widely employed interpreted language, Python faces performance challenges in execution efficiency. Just-in-time (JIT) compilers have been introduced to the Python ecosystem to dynamically compile bytecode into machine code, significantly improving program operation speed. However, the complex optimization strategies of JIT compilers may introduce program defects, thereby affecting program stability and reliability. Existing fuzz testing methods for Python interpreters struggle to effectively detect deep optimization defects and non-crashing defects in JIT compilers. To this end, this study proposes PjitFuzz, a coverage-guided defect detection method for Python JIT compilers. First, PjitFuzz proposes five mutation rules based on JIT optimization strategies to generate program variants that trigger the optimization strategies of Python JIT compilers. Second, a coverage-guided dynamic mutation rule selection method is designed to integrate the advantages of different mutation rules and generate diverse program variants. Third, a checksum-based code block insertion strategy is developed to effectively record changes in variable values during program execution and detect inconsistency in the output. Finally, differential testing is performed by combining different JIT compilation options to effectively detect defects in Python JIT compilers. This study compares PjitFuzz with two state-of-the-art Python interpreter fuzzing methods, FcFuzzer and IFuzzer. The experimental results show that PjitFuzz improves defect detection capability by 150% and 66.7% respectively, and outperforms existing methods in terms of code coverage by 28.23% and 15.68% respectively. For the validity rate of generated test programs, PjitFuzz outperforms the comparative methods by 42.42% and 62.74% respectively. In an eight-month experiment, PjitFuzz has discovered and reported 16 defects, 12 of which have been confirmed by developers.