← Founder Notes
Archive

Coding benchmarks overrate code. kimi k3, evaluated september 21, hits 95.10% on the swe-bench…

Yethikrishna ROriginal on Threads

coding benchmarks overrate code. kimi k3, evaluated september 21, hits 95.10% on the swe-bench verified subset yet lands eighth of 43 models on the vals index with 57.81%.

the gap between a coding subset and a real task is the story.

Context

Vals AI's Kimi K3 page, with a Vals Index evaluation post dated 16 July 2026, says Kimi K3 scores 57.81 percent on the Vals Index (plus or minus 1.06), number 8 of 43 models, with component results of 95.10 percent on the SWE-bench Verified Vals Index subset, 91.27 percent on the Vibe Code Bench subset and 80.90 percent across three trials of Terminal-Bench 2.1, at temperature 1, up to 262k output tokens and max-effort reasoning. A third-party aggregator shows 57.81 as a Vals Index v2.1 entry.

How it compares

The figures 95.10, 57.81 and number 8 of 43 are first-party from the evaluator, as of the page's evaluation. Evaluated September 21 is not supported, since the page dates the evaluation to 16 July 2026. It is a Vals Index subset of SWE-bench Verified and not the full benchmark. Coding benchmarks overrate code is the author's inference: the Vals Index is a composite with non-coding components such as finance tasks, so a lower composite rank does not show coding scores are overstated, and the same page reports high coding components. A ranking among 43 models at one date is not a matched coding-only comparison.

Watch next

  • Vals' current Kimi K3 listing and index version.

Sources

  1. Kimi K3 (Vals AI)vals.ai

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 21 September 2026 at 11:17 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/coding-benchmarks-overrate-code-kimi-k3-evaluated-september-Ddicb9vDOLq" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Coding benchmarks overrate code. kimi k3, evaluated september 21, hits 95.10% on the swe-bench…"></iframe>

More notes