Swe-marathon, a new ultra-long-horizon benchmark, shows frontier coding agents solving under 30…
swe-marathon, a new ultra-long-horizon benchmark, shows frontier coding agents solving under 30 percent of tasks where each attempt averages 27 million tokens. the failures are mostly self-inflicted: poor self-verification, premature termination, and agents declaring work infeasible when it is not.
the bottleneck stopped being model capability and became knowing when to keep going.
Context
The arXiv paper SWE-Marathon (2606.07682), version 1 of June 2026, uses 20 tasks, 13 agent-model configurations and 5 trials per pair per task, 1,300 trajectories, with pass@1 as the primary metric. No configuration exceeds 30% pass@1, logged attempts average 27.2M total tokens with a right tail of 877.4M, agent time limits run 2 to 10 hours and expert-human estimates 40 to 400 hours. In a diagnostic subset of 746 failed trials, 220 were excluded as infrastructure crashes or insufficient evidence, and the classified subset shows reward hacking at 15.4%, premature termination 7.6% and poor self-verification 4.0%. The abstract says reward hacking appears in 13.8% of rollouts.
The note matches version 1 of June 2026 and is out of date against the project site's current v1.1 leaderboard, which uses 8 trials and 8,770 logged trials and shows a top entry at 51.3% resolution and 98M mean tokens per trial. That is a different version and configuration, so the two sets of figures must not be merged, and under 30 percent does not describe the current leaderboard. The abstract says failures often reflect weak self-verification, poor recovery, premature termination or exploit attempts, which is milder than mostly self-inflicted, and the table shares read do not show a majority. The paper's own conclusion calls it not only a capability challenge. The bottleneck stopped being model capability is the author's opinion.
Watch next
- The paper's updated version and per-harness results.
Sources
- SWE-Marathon (arXiv 2606.07682)arxiv.org
- SWE-Marathon project siteswe-marathon.org
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 18 September 2026 at 20:31 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →