← Founder Notes
Archive

Swe-marathon, a new ultra-long-horizon benchmark, shows frontier coding agents solving under 30…

Yethikrishna ROriginal on Threads

swe-marathon, a new ultra-long-horizon benchmark, shows frontier coding agents solving under 30 percent of tasks where each attempt averages 27 million tokens. the failures are mostly self-inflicted: poor self-verification, premature termination, and agents declaring work infeasible when it is not.

the bottleneck stopped being model capability and became knowing when to keep going.

Context

The arXiv paper SWE-Marathon (2606.07682), version 1 of June 2026, uses 20 tasks, 13 agent-model configurations and 5 trials per pair per task, 1,300 trajectories, with pass@1 as the primary metric. No configuration exceeds 30% pass@1, logged attempts average 27.2M total tokens with a right tail of 877.4M, agent time limits run 2 to 10 hours and expert-human estimates 40 to 400 hours. In a diagnostic subset of 746 failed trials, 220 were excluded as infrastructure crashes or insufficient evidence, and the classified subset shows reward hacking at 15.4%, premature termination 7.6% and poor self-verification 4.0%. The abstract says reward hacking appears in 13.8% of rollouts.

How it compares

The note matches version 1 of June 2026 and is out of date against the project site's current v1.1 leaderboard, which uses 8 trials and 8,770 logged trials and shows a top entry at 51.3% resolution and 98M mean tokens per trial. That is a different version and configuration, so the two sets of figures must not be merged, and under 30 percent does not describe the current leaderboard. The abstract says failures often reflect weak self-verification, poor recovery, premature termination or exploit attempts, which is milder than mostly self-inflicted, and the table shares read do not show a majority. The paper's own conclusion calls it not only a capability challenge. The bottleneck stopped being model capability is the author's opinion.

Watch next

  • The paper's updated version and per-harness results.

Sources

  1. SWE-Marathon (arXiv 2606.07682)arxiv.org
  2. SWE-Marathon project siteswe-marathon.org

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 18 September 2026 at 20:31 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/swe-marathon-a-new-ultra-long-horizon-benchmark-Ddbtb-UAjnd" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Swe-marathon, a new ultra-long-horizon benchmark, shows frontier coding agents solving under 30…"></iframe>

More notes