Source note

An AI system to help scientists write expert-level empirical software

Software Foundation ModelCode IntelligenceAutomated Software ProductionGenerative EngineeringAI For Science

ERA is an LLM plus tree-search system that writes scientific software for empirical tasks and optimizes a task-specific quality metric. The excerpt claims it beats strong human or institutional baselines across bioinformatics, epidemiology, and other scientific domains.

  • Scientific discovery can slow down when researchers must write custom software for computational experiments by hand.
  • Many empirical tasks have large design spaces, so one-pass code generation can miss better algorithms or model variants.
  • Better automated software generation matters because it can test more ideas and reduce the time needed to build experiment-specific tools.
  • ERA uses an LLM to propose, implement, and revise scientific software.
  • Tree search keeps multiple candidate solution paths, scores them with a quality metric, and expands the strongest branches.
  • The system can read and use external research ideas, then test whether the resulting code improves the metric.
  • It treats software creation as iterative optimization: generate code, evaluate it, keep strong variants, and search again.
  • Bioinformatics: ERA discovered 40 novel methods for single-cell data analysis that outperformed the top human-developed methods on a public leaderboard.
  • Epidemiology: ERA generated 14 COVID-19 hospitalization forecasting models that outperformed the CDC ensemble and all other individual models in the reported benchmark.
  • Other domains: the excerpt says ERA produced expert-level software in 4 additional areas: geospatial analysis, zebrafish neural activity prediction, numerical integration, and time-series forecasting.
  • The excerpt reports no aggregate metric values, error rates, or leaderboard scores beyond the counts of 40 methods and 14 models plus the stated baseline wins.