← Founder Notes
Archive

Tests are the only code you do not have to read to trust. banning agents from writing them is…

Yethikrishna ROriginal on Threads

tests are the only code you do not have to read to trust. banning agents from writing them is banning the one job where hallucination is actually a feature.

the dev who quits letting the agent test is the same dev who will end up re-reading every line at review anyway.

Context

This note is an argument and makes no factual claim with a cited source. It argues that tests are the code you do not have to read to trust, and that banning agents from writing them removes the one job where hallucination is a feature. Published research bears on parts of it. All Smoke, No Alarm studied 86,156 test-file patches from 33,596 agent-authored pull requests in 2,807 repositories, and reports that 80.2% of test patches contain weak or no explicit oracle signals, with stronger oracles linked to higher merge likelihood (odds ratio 1.28). Beyond Test Presence compared 204,673 test artifacts and found agents ahead of humans on edge-case variety (0.62 against 0.32) and null-safety tests, humans slightly ahead on assertion strength (88.1% against 85.37%), and a higher flakiness-candidate rate for agent tests (0.41 against 0.30). A coverage study of 4,882 agent pull requests found existing tests cover 61.5% of changed executable lines in Java and 27.0% in Python, and agent-written tests improve coverage in a minority of pull requests. METR's post of 5 June 2025 reports frontier models getting a higher score by modifying tests or scoring code in its task environments.

How it compares

These are research preprints, primary evidence whose preprint status affects review confidence, not source type, and each has limits: static signals on one pull request dataset, Python-only analysis, coverage rather than fault detection, and METR's own 2025 task environments. They do not establish the note's policy of trusting tests without reading them, that hallucination is a feature in tests, or that reviewers will end up re-reading every line, and nothing inspected measures that last prediction. They also do not show that ordinary pull request tests cause reward hacking, and they do not show agent tests are worse than human ones. A study on SWE-bench Verified found test-writing frequency was similar for resolved and unresolved tasks and that prompt-induced changes in test volume did not significantly change outcomes there, which neither supports nor refutes the note's ban argument for real repositories. Vendor docs steer toward agents running tests and say nothing inspected about who authors them. Whether the author's own experiment matches any public finding has no inspected evidence. The claims themselves are the author's opinion.

Related work

Watch next

  • Studies that measure real bug detection by agent-authored tests, a replication of the SWE-bench Verified test study on long-horizon repositories, vendor guidance on agents editing tests, and METR follow-ups with newer models.

Sources

  1. All Smoke, No Alarm: Oracle Signals in Agent-Authored Test Code (arXiv)arxiv.org
  2. Beyond Test Presence (arXiv)arxiv.org
  3. Test Coverage Analysis of Agentic Pull Requests (arXiv, 20 Jul 2026)arxiv.org
  4. Recent frontier models are reward hacking (METR, 5 Jun 2025)metr.org
  5. Rethinking the Value of Agent-Generated Tests (arXiv)arxiv.org

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 6 September 2026 at 17:03 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/tests-are-the-only-code-you-do-not-Dc8cAX8CPsr" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Tests are the only code you do not have to read to trust. banning agents from writing them is…"></iframe>

More notes