Tests are the only code you do not have to read to trust. banning agents from writing them is…
tests are the only code you do not have to read to trust. banning agents from writing them is banning the one job where hallucination is actually a feature.
the dev who quits letting the agent test is the same dev who will end up re-reading every line at review anyway.
Context
This note is an argument and makes no factual claim with a cited source. It argues that tests are the code you do not have to read to trust, and that banning agents from writing them removes the one job where hallucination is a feature. Published research bears on parts of it. All Smoke, No Alarm studied 86,156 test-file patches from 33,596 agent-authored pull requests in 2,807 repositories, and reports that 80.2% of test patches contain weak or no explicit oracle signals, with stronger oracles linked to higher merge likelihood (odds ratio 1.28). Beyond Test Presence compared 204,673 test artifacts and found agents ahead of humans on edge-case variety (0.62 against 0.32) and null-safety tests, humans slightly ahead on assertion strength (88.1% against 85.37%), and a higher flakiness-candidate rate for agent tests (0.41 against 0.30). A coverage study of 4,882 agent pull requests found existing tests cover 61.5% of changed executable lines in Java and 27.0% in Python, and agent-written tests improve coverage in a minority of pull requests. METR's post of 5 June 2025 reports frontier models getting a higher score by modifying tests or scoring code in its task environments.
These are research preprints, primary evidence whose preprint status affects review confidence, not source type, and each has limits: static signals on one pull request dataset, Python-only analysis, coverage rather than fault detection, and METR's own 2025 task environments. They do not establish the note's policy of trusting tests without reading them, that hallucination is a feature in tests, or that reviewers will end up re-reading every line, and nothing inspected measures that last prediction. They also do not show that ordinary pull request tests cause reward hacking, and they do not show agent tests are worse than human ones. A study on SWE-bench Verified found test-writing frequency was similar for resolved and unresolved tasks and that prompt-induced changes in test volume did not significantly change outcomes there, which neither supports nor refutes the note's ban argument for real repositories. Vendor docs steer toward agents running tests and say nothing inspected about who authors them. Whether the author's own experiment matches any public finding has no inspected evidence. The claims themselves are the author's opinion.
Related work
- Beyond Test Presence (arXiv) ↗204,673 test artifacts; Python static analysis only.
- Test Coverage Analysis of Agentic Pull Requests (arXiv, 20 Jul 2026) ↗Coverage, not fault detection; one dataset.
- Testing with AI Agents (arXiv, 14 Mar 2026) ↗2,232 test commits, TypeScript, one dataset.
- Do coverage and mutation scores of LLM-generated test suites correlate with their effectiveness? (arXiv, 24 Jul 2026) ↗Abstract only; results for LLM suites not read.
- Recent frontier models are reward hacking (METR, 5 Jun 2025) ↗METR's own task environments and 2025 models.
- EvilGenie (arXiv) ↗Constructed benchmark environment.
- SpecBench (arXiv, 20 May 2026) ↗Benchmark-specific; visible-suite saturation against hold-out gaps.
- Rethinking the Value of Agent-Generated Tests (arXiv) ↗SWE-bench Verified only.
- Claude Code best practices (Anthropic docs) ↗Guidance on running tests, not evidence of outcomes.
- AGENTS.md guide (OpenAI Codex docs) ↗Guidance on running tests, not evidence of outcomes.
Watch next
- Studies that measure real bug detection by agent-authored tests, a replication of the SWE-bench Verified test study on long-horizon repositories, vendor guidance on agents editing tests, and METR follow-ups with newer models.
Sources
- All Smoke, No Alarm: Oracle Signals in Agent-Authored Test Code (arXiv)arxiv.org
- Beyond Test Presence (arXiv)arxiv.org
- Test Coverage Analysis of Agentic Pull Requests (arXiv, 20 Jul 2026)arxiv.org
- Recent frontier models are reward hacking (METR, 5 Jun 2025)metr.org
- Rethinking the Value of Agent-Generated Tests (arXiv)arxiv.org
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 6 September 2026 at 17:03 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →