← Founder Notes
Archive

A new benchmark called fixedbench gives coding agents repos where the fix is already applied and…

Yethikrishna ROriginal on Threads

a new benchmark called fixedbench gives coding agents repos where the fix is already applied and the correct answer is an empty patch, and five frontier models across four harnesses keep failing to abstain. the paper from september 16 shows agents cannot tell already-done from to-do.

knowing when to do nothing is the skill that still separates them from engineers.

Context

The paper Coding Agents Don't Know When to Act, by Gloaguen, Mundler, Muller, Raychev and Vechev of ETH Zurich and LogicStar.ai, arXiv 2605.07769, builds FixedBench from a random subset of 200 SWE-Bench Verified instances with the fix already applied, where the correct behavior is an empty patch. It tests four closed models and one open-source model with their own agent harnesses, and reports unnecessary changes in 35-65% of instances, shifted by prompt: for GPT-5.4 mini abstention was 60.5% with one prompt, 36.5% when told to edit and 88.5% with a reproduce-then-abstain prompt.

How it compares

The authors' benchmark and results are their own. Keep failing to abstain overstates it: they fail in a large share, not all, and a prompt largely reduces it with a different over-abstention failure mode. The five models are four closed and one open-source, one of them a mini, with harness mapping that is approximate. The paper read is version 1 with a May 2026 identifier and the OpenReview entry is dated 23 May 2026; no source dated 16 September was found. Scores are not merged across prompts or harnesses. The line about knowing when to do nothing is the author's opinion.

Watch next

  • The promised code release and results with newer models.

Sources

  1. Coding Agents Don't Know When to Act (arXiv 2605.07769)arxiv.org

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 19 September 2026 at 02:46 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/a-new-benchmark-called-fixedbench-gives-coding-agents-DdcYSm_DINP" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="A new benchmark called fixedbench gives coding agents repos where the fix is already applied and…"></iframe>

More notes