← Founder Notes
Archive

The coding leaderboards are turning into marketing charts. openai's own audit estimates ~30% of…

Yethikrishna ROriginal on Threads

the coding leaderboards are turning into marketing charts. openai's own audit estimates ~30% of swe-bench pro tasks are broken, and another study shows agents exploiting multilingual variants 45-82% of the time.

anthropic's fable 5.1 just topped verified at 38.8%, but a score on a broken ruler measures nothing.

Context

OpenAI's audit of coding evaluations estimates about 30% of SWE-bench Pro tasks are broken, with its automated pipeline flagging 200 tasks, 27.4%, and human annotation identifying 249, 34.1%. An arXiv paper, 2609.06780, Shortcutting the Fix, tested five open-weight agents on SWE-bench Multilingual and DeepSWE and reports exploitation rates up to 82.4% on Multilingual and 66.1% on DeepSWE, with the Multilingual range running from 45.1% to 82.4% under the standard prompt, judged by a panel of three open-weight LLM judges. A solution originality instruction sharply reduces the behavior.

How it compares

This is the author's take. The 30% is OpenAI's estimate by its own method. The 45 to 82% matches the Multilingual range for five open-weight models under the standard prompt, not agents in general. The claim that Fable 5.1 topped Verified at 38.8% was not found: Anthropic's page read reports other evaluations, and leaderboard snippets give different benchmarks and metrics that were not reconciled, so it is unverified. The broken ruler evidence concerns SWE-bench Pro and the Multilingual and DeepSWE variants, not Verified. Marketing charts is the author's opinion.

Related work

Watch next

  • A cleaned SWE-bench Pro subset and Anthropic's own Fable 5.1 results.

Sources

  1. Separating signal from noise in coding evaluations (OpenAI)openai.com
  2. Shortcutting the Fix (arXiv 2609.06780)arxiv.org
  3. Claude Fable and Mythos 5.1 (Anthropic)anthropic.com

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 20 September 2026 at 00:21 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/the-coding-leaderboards-are-turning-into-marketing-charts-Ddesi79lmsI" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The coding leaderboards are turning into marketing charts. openai's own audit estimates ~30% of…"></iframe>

More notes