The boring part of compilers is where agents actually shine. a september 16 paper, peepholebench,…
the boring part of compilers is where agents actually shine. a september 16 paper, peepholebench, turns real missed llvm instcombine optimizations — the small localized patches llvm's own pipeline never catches — into a coding-agent test.
the unglamorous 20-line fixes are the benchmark now.
Context
PeepholeBench is from arXiv 2607.02684, Can Coding Agents Implement Missed Compiler Optimizations?, by University of Waterloo authors. It derives tasks from 21 resolved LLVM issues and 19 merged pull requests, gives agents only the issue context that existed before each fix, and scores patches for correctness and profitability against human fixes. It tested Gemini CLI with Gemini 3 Flash, Codex CLI with GPT 5.4 and GPT 5.4 mini, and Claude Code with Sonnet 4.6. The abstract reports a tension between correctness and profitability and that no agent matches human developers on both at once.
The note says a September 16 paper. The arXiv ID and search date are July 2026, so the September date is not supported. The headline finding is that no agent matched humans on both dimensions, with overly narrow transformations and misuse of LLVM-specific mechanisms as the dominant failures. That does not support the opening line that agents shine here, which is the author's take. The task set is small at 21 issues, and per-agent results were not extracted.
Watch next
- The per-agent results table and the task list.
Sources
- arXiv 2607.02684 full textarxiv.org
- arXiv 2607.02684 abstractarxiv.org
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 21 September 2026 at 20:39 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →