← Founder Notes
Archive

A new real-swe benchmark ran eight frontier coding models against licensed production codebases

Yethikrishna ROriginal on Threads

a new real-swe benchmark ran eight frontier coding models against licensed production codebases: the best scored 38.8%, and seven of eight didn't clear a third of tasks. swe-bench says 96%, real code says otherwise.

the gap between leaderboard and production is the entire game now.

Context

Specific's Real-SWE benchmark pages say tasks come from private production codebases licensed from a real company, with resolution rate as pass@1 averaged over eight independent runs per task with 95% confidence intervals, using native harnesses, high reasoning for all models, isolated sandboxes and verifiers injected at grading. A secondary report dated 14 September 2026 says the top score was 38.8% for Fable 5.1, with Gemini 3.8 Flash 7.6 points behind and ten tickets.

How it compares

This is vendor-reported by Specific, not independent. The first-party pages now show a different leaderboard: GPT-6 Astra (Codex CLI) 46.25%, Fable 5.1 (Claude Code) 45.00%, Gemini 3.8 Flash (Gemini CLI) 38.75%, GLM 5.3 37.50%, Muse Spark 1.3 36.25%, Grok 4.6 32.50%, GPT-5.6 Sol 26.25% and Kimi K3 20.00%, eight model-harness pairs. The 38.8% top score matches an earlier snapshot reported on 14 September, not the current top of 46.25%, and seven of eight below a third holds only for that earlier snapshot, not today's table. The first-party task count was not inspected, and SWE-bench says 96% has no inspected source.

Related work

Watch next

  • A dated leaderboard history.

Sources

  1. Real-SWE (Specific)withspecific.com
  2. Real-SWErealswe.com

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 18 September 2026 at 10:35 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/a-new-real-swe-benchmark-ran-eight-frontier-DdapLJmCMUX" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="A new real-swe benchmark ran eight frontier coding models against licensed production codebases"></iframe>

More notes