The real codebases just humbled every coding agent. real-swe, out september 14, ran 640 rollouts on…
the real codebases just humbled every coding agent. real-swe, out september 14, ran 640 rollouts on licensed production code and the best model closed 38.8 percent of tasks at 6.96 dollars each.
the benchmark gap is the pricing gap.
Context
Specific Labs' Real-SWE page, dated September 2026, says tasks come from a private production codebase licensed from a real-world company, that resolution rate is equivalent to pass@1 averaged over eight independent runs per task with 95 percent confidence intervals, and that models use native harnesses such as Codex CLI and Claude Code at high reasoning. A secondary article of September 14, 2026 says Specific Labs published it on September 12, ran eight models against ten licensed private-codebase tickets, and had Fable 5.1 via Claude Code lead at 38.8 percent at 6.96 dollars a rollout, with one task at 0 for 64.
The 640 rollouts figure is consistent with 10 tasks, 8 models and 8 runs, but it was not read in first-party text. The September 14 date is the secondary article date. The current first-party snapshot lists Fable 5.1 on Claude Code at 45.00 percent and GPT-6 Astra on Codex CLI at 46.25 percent as first, so 38.8 percent and 6.96 dollars belong to an earlier window, and the two runs are different and not contradictory. The scores are vendor-reported, the private repositories limit public reproducibility, and no outside verification was established. Rollout cost estimates range from 2.50 to 6.96 dollars per the snapshot, which does not itself establish a pricing gap. The benchmark gap is the pricing gap is the author's take.
Related work
- the best coding model on real enterprise code still fails most tasks, ↗An earlier note on the same benchmark.
- a new real-swe benchmark ran eight frontier coding models against lice ↗An earlier note on the same benchmark.
Watch next
- Leaderboard changes and any independent replication.
Sources
- Specific Labs: Real-SWE benchmarkwithspecific.com
- AI Model Report: Real-SWE benchmark (September 14, 2026)aimodelreport.com
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 23 September 2026 at 02:49 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →