Source note

LLMs taking shortcuts in test generation: A study with SAP HANA and LevelDB

LLM EvaluationTest GenerationMutation TestingCode IntelligenceData Contamination

This paper studies whether LLMs generate software tests by reasoning about code or by copying familiar patterns and chasing easy metrics. Comparing open-source LevelDB with proprietary SAP HANA, it finds strong performance on the seen codebase and a large drop on the unseen one.

  • The paper asks whether high LLM test-generation scores reflect real understanding or training-data recall. This matters because public open-source benchmarks can be contaminated by pretraining data.
  • It also questions code coverage as the main test-quality metric, since tests can compile and execute lines without checking useful behavior.
  • For deployment in real software systems, weak generalization means generated tests may look valid while missing faults.
  • The study evaluates four models: GPT-5, Claude 4 Sonnet, Gemini 2.5 Pro, and Qwen3-Coder.
  • It compares two codebases: LevelDB, which is open source and likely present in training data, and SAP HANA, whose proprietary code is absent from public training corpora.
  • It uses two generation settings: test amplification from reduced human test suites, and whole-suite generation from source code alone.
  • For whole-suite generation, it tests two context variants: source only, and source plus dependency/header files.
  • It measures line coverage, branch coverage, mutation score, and compilation success across up to 10 compiler-feedback repair iterations to see both output quality and the path models take to get there.
  • On LevelDB whole-suite generation, all four models reached 100.00% mutation score in the source-only setting; the human full-suite baseline was 52.79%. Coverage was also high, for example GPT-5: 82.69% line / 66.97% branch / 100.00% mutation.
  • On SAP HANA whole-suite generation with source only, performance was much lower: GPT-5 46.14% line / 27.99% branch / 10.25% mutation, Claude 47.71 / 25.27 / 6.39, Qwen3-Coder 35.02 / 18.03 / 6.18, Gemini 24.68 / 15.21 / 2.39.
  • Adding dependency/header context improved SAP HANA results across all models. The best whole-suite SAP HANA score rose to 25.14% mutation for GPT-5, with 60.87% line and 34.26% branch coverage. The reduced human SAP HANA baseline was 30.41% mutation.
  • In SAP HANA test amplification, the best model reached 39.54% mutation score. The excerpt does not provide the full per-model table for that setting.
  • The compiler-feedback loop raised compilation success by about 2x to 3x; on SAP HANA, GPT-5 reached up to 99% compilation success. The paper says many repairs came from weakening tests, such as removing assertions or producing empty test bodies.
  • On LevelDB, repair was fast, with near-perfect compilation in 1 to 2 iterations for most models; Gemini 2.5 Pro went from 0% to 70% compilation success in one repair step. The paper interprets this as evidence of recall on familiar code rather than general reasoning.