← Founder Notes
Archive

Kimi k3 ranks second on the agentic benchmark but costs more than opus 4.8 to run and averages…

Yethikrishna ROriginal on Threads

kimi k3 ranks second on the agentic benchmark but costs more than opus 4.8 to run and averages nearly an hour per task. frontier quality is now measured in dollars and hours per task, not accuracy.

the leaderboard that matters is your invoice.

Context

Artificial Analysis's article of 21 July 2026 reports AA-Briefcase, its proprietary agentic knowledge-work benchmark: Kimi K3 scores Elo 1543, second to Claude Fable 5 at 1574, ahead of GPT-5.6 Sol (max) at 1501, Claude Sonnet 5 (max) at 1388 and Claude Opus 4.8 (max) at 1347. Kimi K3 costs 10.57 dollars per task, which AA says is more than Opus 4.8, and averages 56.4 minutes, 83 turns and 120k output tokens per task.

How it compares

Second is as of AA's 21 July article, two months before the note, and the current leaderboard was not checked. The figures are one operator's harness and config and are not merged with other benchmarks. The exact Opus 4.8 per-task cost was not extracted.

Related work

Watch next

  • The live AA-Briefcase page for later rankings.

Sources

  1. Kimi K3 agentic knowledge benchmark (Artificial Analysis, 21 Jul 2026)artificialanalysis.ai

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 18 September 2026 at 04:31 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/kimi-k3-ranks-second-on-the-agentic-benchmark-DdZ_lpMAr7z" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Kimi k3 ranks second on the agentic benchmark but costs more than opus 4.8 to run and averages…"></iframe>

More notes