Benchmarks just moved from answering questions to doing ai research. vals sage, updated september…
benchmarks just moved from answering questions to doing ai research. vals sage, updated september 10, grades 81 models on autonomous llm r&d against published, human and frontier references.
the exam is now can a model build the next model.
Context
Vals AI's benchmark listing has two entries relevant to this note. The page named SAGE is Student Assessment with Generative Evaluation, which grades handwritten math work, was updated September 1, 2026 and shows 23 models, with Claude Opus 4.7 leading at 56.10 percent. A separate Vals RSI Index page, updated August 12, 2026 with 5 tasks and 3 named model harnesses, asks whether a model can do the research that builds the next model, and the benchmarks listing describes it as autonomous LLM R&D scored against published, human and frontier-model references.
The note names SAGE but its description matches the RSI Index listing, and which one the author meant is not asserted here. The 81 models and the September 10 date were not found on either page read, so those two figures are unsupported, not refuted. Scores must not be merged across Vals benchmarks or across harnesses. The exam is now can a model build the next model is the author's take.
Watch next
- Whether the RSI Index updates with a larger model set, and the source of the 81 models.
Sources
- Vals AI: SAGEvals.ai
- Vals AI: RSI Indexvals.ai
- Vals AI: benchmarksvals.ai
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 22 September 2026 at 14:37 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →