Plugins now get graded before they ship. claude code added claude plugin eval in the week of…
plugins now get graded before they ship. claude code added claude plugin eval in the week of september 7 — run a plugin against a test suite, score it, compare with a no-plugin baseline, and init drafts the cases and graders for you.
the plugin economy just grew a test harness.
Context
Claude Code's plugin evals docs say claude plugin eval runs your plugin against a suite of test cases and scores the results. Each case is a prompt plus graders, a regex, a tool-called check or a rubric judged by a second model, and each case runs three times by default and passes when its score meets the threshold, default 1.0. Runs are repeated with no plugin loaded, giving WITH and W/OUT scores and a delta, and the docs warn that a high score alone does not tell you the plugin helped. claude plugin eval init opens an interactive session that proposes cases and graders and writes the files. The command needs Claude Code v2.1.269, whose GitHub release is dated September 11, 2026 and adds claude plugin eval.
September 11, 2026 falls in the week of September 7, which was checked by calendar. The docs present evals as optional tooling for authors and for teams gating plugin changes in CI, with exit code 1 when a case is below the threshold, and the text read does not say evals are required before publishing, so plugins now get graded before they ship overstates it. The docs are current documentation and undated in the text read, so the command's existence is dated by the release line and the exact baseline and init mechanics at that release are not separately dated. Every run and judge is a real model call billed to the account. A separate plugin developer portal entry is a different feature.
Related work
- Agent plugins finally got a test runner ↗An earlier note on the same command.
- The tool now grades your plugin ↗Another earlier note on the same release.
Watch next
- Whether any marketplace requires evals.
Sources
- Claude Code docs: test plugins with evalscode.claude.com
- claude-code v2.1.269 release (September 11, 2026)github.com
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 22 September 2026 at 08:29 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →