Source note

GapForge: Directed Compiler Fuzzing via Coverage-Gap Analysis

Compiler FuzzingCode IntelligenceLLM Test GenerationCoverage Guided TestingAutomated Software Production

GapForge is an LLM-based compiler fuzzing technique that targets specific source-code coverage gaps instead of generating tests without compiler-wide guidance. On GCC 14.3.0 and LLVM 19.1.0, it improves 72-hour coverage and finds 12 real-world compiler failures.

  • Large compilers such as GCC and LLVM contain long-tail, under-tested code regions that general-purpose fuzzers repeatedly miss.
  • Reaching an uncovered region often requires both a particular program structure and specific compiler options, which file-level summaries and program-driven generation do not reliably infer.
  • This matters because compiler defects can cause crashes and silent miscompilations in downstream software.
  • GapForge scores source files with L_f × (1-C_f)^2, favoring large files with low line coverage, then samples one target file per iteration.
  • An LLM analyzes each uncovered line span together with nearby covered context to infer the control-flow, data, and compilation-option requirements for reaching the target basic blocks.
  • It synthesizes a prompt combining general C/C++ generation constraints, these target requirements, and previously failed prompts for the same file.
  • Generated programs are compiled, measured with coverage feedback, and used to guide subsequent target selection and prompt refinement.
  • Within 72 hours, GapForge reached 68.13% line coverage on GCC core modules and 69.11% on LLVM core modules.
  • It covered 24,736 more lines than WhiteFox on GCC and 19,798 more lines on LLVM; WhiteFox reached 64.62% and 65.02%, respectively.
  • Compared with the strongest reported baseline, LegoFuzz, which reached 64.99% on GCC and 66.59% on LLVM, GapForge also improved coverage of the official test suites by 3,452 GCC lines and 531 LLVM lines, versus LegoFuzz's 705 and 143 lines.
  • GapForge found 12 real-world compiler failures: 5 in GCC and 7 in LLVM, including 8 crashes and 4 miscompilations.
  • Nine ablation variants performed worse than the complete system, supporting contributions from target selection, targeted summarization, compilation-option inference, and failure reflection.