A new benchmark called fixedbench gives coding agents repos where the fix is already applied and…
a new benchmark called fixedbench gives coding agents repos where the fix is already applied and the correct answer is an empty patch, and five frontier models across four harnesses keep failing to abstain. the paper from september 16 shows agents cannot tell already-done from to-do.
knowing when to do nothing is the skill that still separates them from engineers.
Context
The paper Coding Agents Don't Know When to Act, by Gloaguen, Mundler, Muller, Raychev and Vechev of ETH Zurich and LogicStar.ai, arXiv 2605.07769, builds FixedBench from a random subset of 200 SWE-Bench Verified instances with the fix already applied, where the correct behavior is an empty patch. It tests four closed models and one open-source model with their own agent harnesses, and reports unnecessary changes in 35-65% of instances, shifted by prompt: for GPT-5.4 mini abstention was 60.5% with one prompt, 36.5% when told to edit and 88.5% with a reproduce-then-abstain prompt.
The authors' benchmark and results are their own. Keep failing to abstain overstates it: they fail in a large share, not all, and a prompt largely reduces it with a different over-abstention failure mode. The five models are four closed and one open-source, one of them a mini, with harness mapping that is approximate. The paper read is version 1 with a May 2026 identifier and the OpenReview entry is dated 23 May 2026; no source dated 16 September was found. Scores are not merged across prompts or harnesses. The line about knowing when to do nothing is the author's opinion.
Watch next
- The promised code release and results with newer models.
Sources
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 19 September 2026 at 02:46 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →