← Founder Notes
Archive

The code review benchmark now has a clear leader. augment's review agent, powered by gpt-5.2 and…

Yethikrishna ROriginal on Threads

the code review benchmark now has a clear leader. augment's review agent, powered by gpt-5.2 and out september 17, beat cursor bugbot and coderabbit by about 10 points on the only public benchmark for ai-assisted review.

reviewing well is now a measured capability.

Context

Augment's blog on why GPT-5.2 is its model of choice for Augment Code Review, dated 11 December 2025 and last updated 18 June 2026, says it has the highest accuracy on the only public benchmark for AI-assisted code review, outperforming systems from Cursor Bugbot, CodeRabbit and others by about 10 points on overall quality. The Martian code-review-benchmark README describes an open replication of the benchmark used by Augment and Greptile: 50 pull requests across 5 codebases, human golden comments and an LLM judge, with judge variance and sparse nit coverage noted as limits.

How it compares

The roughly 10 point gap is a vendor statement on one benchmark with one scoring setup; the page read gives no numbers, benchmark name or judge. Out September 17 is not supported, since the post dates to December 2025 with a June 2026 update. The only public benchmark is Augment's wording; secondary posts describe other benchmarks with other winners, which are different benchmarks and not refutation. The replication README gives no score that was read. Reviewing well is now a measured capability is the author's take.

Watch next

  • Augment's full benchmark analysis and an independent run of the replication.

Sources

  1. Why GPT-5.2 is our model of choice for Augment Code Review (Augment)augmentcode.com
  2. code-review-benchmark offline README (Martian)github.com

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 20 September 2026 at 21:20 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/the-code-review-benchmark-now-has-a-clear-Ddg8nn_jSqt" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The code review benchmark now has a clear leader. augment's review agent, powered by gpt-5.2 and…"></iframe>

More notes