ToxiHeap: LLM-guided JavaScript Engine Fuzzing Framework with Toxic Labeling
Author:
Affiliation:

Clc Number:

TP311

Fund Project:

  • Article
  • |
  • Figures
  • |
  • Metrics
  • |
  • Reference
  • |
  • Related
  • |
  • Cited by
  • |
  • Materials
  • |
  • Comments
    Abstract:

    Existing fuzzer for browser JavaScript engines still have limitations in detecting potential vulnerabilities, especially in identifying non-crashing heap memory errors. This study proposes ToxiHeap, a JavaScript engine fuzzing framework that integrates runtime monitoring based on memory toxic labeling with large language model (LLM)-guided semantic generation as two parallel core components. The detection track instruments heap allocation and access paths in the target engine. Shadow memory and fine-grained toxic labeling provide a unified mechanism for representing multiple heap errors, including use-after-free (UAF), double free, heap out-of-bounds reads and writes (OOB-R/W), memory leak, and uninitialized memory read (UMR). Accesses to freed memory and other illegal states are converted into consumable exception signals, which compensates for the lack of effective detection signals for non-crashing vulnerabilities in approaches that rely solely on crash signals. The generation track employs a two-stage “distill-reward” LLM-based mutator that extracts semantic features from historical proofs of concept (PoCs) and generates state-dependent test cases in a targeted manner, thus overcoming the limitations of purely syntax-based or shallow-semantic mutations in exploring deep execution paths. The two tracks are integrated through a feedback loop driven by coverage gains and non-crashing detection signals. This feedback loop jointly drives sample retention, path weight updates, and knowledge base growth. An isolated replay verification mechanism is also introduced to revalidate anomalies in a clean process, which significantly reduces nondeterministic false positives. Experimental results show that on the JavaScriptCore (JSC) engine, ToxiHeap increases branch coverage from 31.17% to 33.52% in a 24-hour run, and achieves the highest or near-highest branch coverage on other mainstream engines including V8, SpiderMonkey (SM), and ChakraCore (CH). The proportion of valid samples remains above 91%. On a set of 50 PoC references covering multiple vulnerability patterns including UAF, the average detection rate across the four engines reaches 89.18%.

    Reference
    Related
    Cited by
Get Citation

沙乐天,龙章伯,丁加宇,黄海平,肖甫. ToxiHeap: LLM制导和毒性标记的JavaScript引擎模糊测试框架.软件学报,,():1-27

Copy
Share
Article Metrics
  • Abstract:
  • PDF:
  • HTML:
  • Cited by:
History
  • Received:September 02,2025
  • Revised:November 03,2025
  • Adopted:
  • Online: May 13,2026
  • Published:
You are the firstVisitors
Copyright: Institute of Software, Chinese Academy of Sciences Beijing ICP No. 05046678-4
Address:4# South Fourth Street, Zhong Guan Cun, Beijing 100190,Postal Code:100190
Phone:010-62562563 Fax:010-62562533 Email:jos@iscas.ac.cn
Technical Support:Beijing Qinyun Technology Development Co., Ltd.

Beijing Public Network Security No. 11040202500063