← Founder Notes
Archive

The frontier benchmark race has narrowed to a three-point spread. as of september 20, claude fable…

Yethikrishna ROriginal on Threads

the frontier benchmark race has narrowed to a three-point spread. as of september 20, claude fable 5.1 leads benchlm at 84.74, gpt-6 astra sits at 82.81, and claude opus 5 trails at 81.87.

the models are converging faster than the marketing.

Context

BenchLM's overall ranking page, last updated 2 October 2026, shows a composite across 8 benchmark categories with provisional and verified ranks and conditional range columns, and says ranges do not establish confidence in a rank. On that snapshot GPT-6 Astra is 88.8 (rank 1), Claude Opus 5.5 87.8, Claude Sonnet 5.5 83.4, Claude Fable 5.1 82.9 (rank 4) and Claude Opus 5 80.0.

How it compares

None of the note's 20 September numbers (84.74, 82.81, 81.87) appear on the 2 October page, and that snapshot cannot establish the earlier ones. Different window, not refutation, and the historical figures are unverified. A three-point spread sits inside the stated ranges, so the leads and narrowed to are not statistically established, and the three-point spread inference should not be applied to the older numbers. The composite is not comparable with vendor-published benchmarks. The models are converging faster than the marketing is the author's take.

Watch next

  • A dated 20 September snapshot before accepting the figures.

Sources

  1. BenchLM best overall (updated 2 Oct 2026)benchlm.ai

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 21 September 2026 at 00:06 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/the-frontier-benchmark-race-has-narrowed-to-a-DdhPoQckesQ" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The frontier benchmark race has narrowed to a three-point spread. as of september 20, claude fable…"></iframe>

More notes