The largest study of ai coding agents by pr acceptance found no best agent — codex wins…
the largest study of ai coding agents by pr acceptance found no best agent — codex wins consistency, claude code wins docs at 92.3%, cursor wins fixes at 80.4%. the best agent question is the wrong one.
you're picking a specialist for the task you have, not a champion.
Context
An arXiv paper accepted at MSR 2026 compares five agents, OpenAI Codex, GitHub Copilot, Devin, Cursor and Claude Code, over 7,156 pull requests filtered from the 33,596-PR AIDev dataset, using repositories with 100 or more stars and a task-stratified comparison. The dataset paper, also MSR 2026 (13 to 14 April 2026), is the larger resource. The study reports task type as a dominant factor, with a 29 percentage point gap, and Codex consistently high across nine task categories, from 59.6% on tests to 88.6% on chores.
Claude Code wins docs at 92.3% was not found in any text read; the study table lists Claude Code with only 139 PRs and 71.9% overall acceptance, so a per-category rate would rest on a very small sample, and the 139 is an overall count, not the docs denominator. A search excerpt shows Cursor 80.4% on fixes against Codex 83.0%, which contradicts Cursor winning fixes, but that excerpt was not located in the fetched page text. Largest is not stated by the study. This is observational data on public repositories with non-random task assignment, so no best agent is the author's framing.
Related work
- AIDev dataset paper (arXiv, MSR 2026) ↗Text inspected.
Watch next
- The peer-reviewed version and per-agent per-category tables.
Sources
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 18 September 2026 at 05:47 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →