Alibaba open-sourced a code review tool that pairs a deterministic pipeline with an llm agent, and…
alibaba open-sourced a code review tool that pairs a deterministic pipeline with an llm agent, and it gained 3,286 stars in a single day on the way past 34,000. its own aacr-bench shows it beating claude code at review.
the part of development everyone considered too boring for ai turned out to be the part where open source wins first.
Context
The alibaba/open-code-review repository, read on 4 October 2026, describes a hybrid code review tool with deterministic pipelines and an LLM agent, line-level comments and a built-in ruleset for issues such as null pointers, thread safety, XSS and SQL injection. It shows an Apache-2.0 license, a creation date of 18 May 2026 and 21,269 stars. Its README says that against general-purpose agents such as Claude Code it has significantly higher precision and F1 with the same underlying model, at about a ninth of the tokens, and that its recall is lower. The benchmark it cites, AACR-Bench, has 50 repositories, 200 pull requests, 10 languages and 1,505 annotated issues.
Beating Claude Code holds on precision and F1 only; the README says recall is lower. The tool and the benchmark are both Alibaba's and the results are vendor-reported, not independent, and exact scores were not read. The 21,269 stars at fetch is below the note's 34,000, so the note's figure and the 3,286 in a day are unverified. Repo creation is not a release date. Open source winning first in the part of development everyone considered too boring is the author's opinion.
Watch next
- Independent code-review evaluations.
Sources
- open-code-review (GitHub)github.com
- aacr-bench (GitHub)github.com
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 19 September 2026 at 00:02 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →