Benchmarks finally stopped grading diffs. backendforge, out since july, runs 56 backend tasks where…
benchmarks finally stopped grading diffs. backendforge, out since july, runs 56 backend tasks where an llm must ship a dockerized service that passes http tests against an openapi contract.
agents pass or fail on shipped behavior.
Context
BackendForge is described in an arXiv paper listed on July 13, 2026 by authors at Peking University and Fudan University. It has 56 contract-defined backend generation tasks rewritten from real open-source applications, with OpenAPI contracts defining 2,345 API operations. Given a visible spec and contract, a model must produce a Dockerized service that is built, deployed and evaluated only through HTTP tests, and scoring does not inspect source code. A base oracle is made first, then a test agent and a code agent co-evolve the test oracle and reference service, with a review agent filtering tests. The authors report the best model solving 55.4% under the base oracle and 28.6% under the final oracle (16 of 56 tasks), and every other configuration at most 3 of 56 under the final oracle.
The results are the authors' own from a preprint, and the base oracle and final oracle are different measures. The paper's limits are Python backend services tested through black-box HTTP only, with no UI, other stacks, performance, deployment security or production hardening. It positions itself against repository-repair benchmarks and itself cites earlier backend benchmarks that grade deployed behavior, BaxBench and ABC-Bench, so finally is the author's framing.
Watch next
- Replication on other stacks.
Sources
- arXiv: BackendForge (July 13, 2026)arxiv.org
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 22 September 2026 at 12:19 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →