A new real-swe benchmark ran eight frontier coding models against licensed production codebases
a new real-swe benchmark ran eight frontier coding models against licensed production codebases: the best scored 38.8%, and seven of eight didn't clear a third of tasks. swe-bench says 96%, real code says otherwise.
the gap between leaderboard and production is the entire game now.
Context
Specific's Real-SWE benchmark pages say tasks come from private production codebases licensed from a real company, with resolution rate as pass@1 averaged over eight independent runs per task with 95% confidence intervals, using native harnesses, high reasoning for all models, isolated sandboxes and verifiers injected at grading. A secondary report dated 14 September 2026 says the top score was 38.8% for Fable 5.1, with Gemini 3.8 Flash 7.6 points behind and ten tickets.
This is vendor-reported by Specific, not independent. The first-party pages now show a different leaderboard: GPT-6 Astra (Codex CLI) 46.25%, Fable 5.1 (Claude Code) 45.00%, Gemini 3.8 Flash (Gemini CLI) 38.75%, GLM 5.3 37.50%, Muse Spark 1.3 36.25%, Grok 4.6 32.50%, GPT-5.6 Sol 26.25% and Kimi K3 20.00%, eight model-harness pairs. The 38.8% top score matches an earlier snapshot reported on 14 September, not the current top of 46.25%, and seven of eight below a third holds only for that earlier snapshot, not today's table. The first-party task count was not inspected, and SWE-bench says 96% has no inspected source.
Related work
- Real-SWE benchmark shows best frontier coding agent solves only 38.8% (aimodelreport, 14 Sep 2026) ↗Secondary; earlier snapshot.
Watch next
- A dated leaderboard history.
Sources
- Real-SWE (Specific)withspecific.com
- Real-SWErealswe.com
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 18 September 2026 at 10:35 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →