The most dependable agent deployments are the boring ones. doordash ran a multi-agent system over…
the most dependable agent deployments are the boring ones. doordash ran a multi-agent system over 60,000 feature flags across 623 repos and got 45 usable pull requests out of 50, at $4.79 and 13.8 minutes each versus one to two hours by hand.
the highest-return agent work right now is deleting code, not writing it.
Context
DoorDash's engineering blog, dated about 24 August 2026 and described as presented at the ICSME 2026 industry track, says its experimentation platform manages over 60,000 feature flags across roughly 623 repositories, and that it identified over a thousand stale flags. A multi-agent system took a Jira ticket to a merge-ready pull request, and for the 50 most recent stale flags it processed, it produced a merge-ready PR for 45, at an average of 4.79 dollars and 13.8 minutes, against the one to two hours a manual cleanup would take. Of the 50, 31 merged on the first shot, 14 needed one minor revision and five needed an engineer; no bugs or regressions were observed.
This is the author's take on a company-reported case. The 60,000 flags and 623 repositories are the platform's scale, and the evaluation covered 50 flags across several Kotlin repositories. Usable and merge-ready are DoorDash's wording, and the one to two hours is DoorDash's estimate, not a measured baseline. One case does not show that the most dependable deployments are the boring ones, and that deleting code is the highest return is the author's opinion.
Watch next
- DoorDash results beyond the 50-flag sample.
Sources
- Automating feature flag cleanup at scale (DoorDash)careersatdoordash.com
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 20 September 2026 at 00:39 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →