← Founder Notes
Archive

99.9% on arc-agi-3, then 62.7% in a different harness. that is not a benchmark bug — it is the eval…

Yethikrishna ROriginal on Threads

99.9% on arc-agi-3, then 62.7% in a different harness. that is not a benchmark bug — it is the eval market splitting in public.

whoever ships the unified harness by q1 owns the whole category.

Context

ARC Prize's results page for GPT-6 Astra (2 September 2026) and its blog (3 September 2026) report two harness configurations of the same model on the ARC-AGI-3 Semi-Private split. With the Standard harness, where the model carries forward notes it chooses to keep, the best observed is 62.7% at max reasoning for about 26K USD. With the Provider Adapter harness, which preserves opaque reasoning state between requests and uses compaction, the best observed is 99.9% at high reasoning for about 19K USD. The leaderboard table lists Provider Adapter High at 99.95%, and Standard Max at 62.71%.

How it compares

These are two configurations, not one score falling. The blog's 99.9% and the table's 99.95% are unreconciled here, and are not called rounding. Each value is kept with its own source: the blog gives 99.9% and the table gives 99.95%. The max and high reasoning levels and the costs belong to their own configurations. OpenAI's own page claims 99.9% on ARC-AGI-3 without naming a harness in the sentence read. Not a benchmark bug, the eval market splitting, and a unified harness owning the category by the first quarter are the author's opinion and prediction.

Related work

Watch next

  • A reconciliation of the blog and leaderboard values.

Sources

  1. GPT-6 Astra results (ARC Prize, 2 Sep 2026)arcprize.org
  2. Astra on ARC-AGI-3 (ARC Prize, 3 Sep 2026)arcprize.org

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 6 September 2026 at 01:02 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/99-9-on-arc-agi-3-then-62-Dc6uFiziPu-" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="99.9% on arc-agi-3, then 62.7% in a different harness. that is not a benchmark bug — it is the eval…"></iframe>

More notes