Executable repositories and multi-attempt memory
Repository-level code generation is getting judged by whether the whole project installs, links, runs, and survives repeated attempts. EnvGraph treats runtime failure diagnosis as a structured attribution problem across external packages and internal references, and reports gains of 5.72 to 5.87 points in functional correctness and 4.58 to 8.66 points in non-functional quality on RAL-Bench and NL2Repo-Bench. LiveCoder attacks the same bottleneck from another angle: it keeps success notes, failure notes, and the best repository artifact across attempts. On RAL-Bench it reports up to +22.94 functional points, up to 81.58% repository reuse, and up to 53.63% cost reduction. The shared message is practical: execution feedback is now part of the method, not just the scorecard.