Openai designated gpt-6 astra as the first model at its critical cybersecurity threshold, meaning…
openai designated gpt-6 astra as the first model at its critical cybersecurity threshold, meaning it can find unknown flaws and build exploits on well-protected systems without step by step guidance, and delayed part of the release to harden it. without safeguards it scored 100 percent on exploitbench and found two previously unknown vulnerabilities during evaluation.
the selling point of frontier models is no longer what they can do, it is what they are stopped from doing.
Context
OpenAI's post of 1 September 2026 says it believes Astra meets the Critical cybersecurity capability threshold under its Preparedness Framework, and that over several weeks it delayed parts of Astra's development and release while strengthening protections. It reports a perfect score of 100% on ExploitBench, and says that on an internal ExploitBench port of 20 recent V8 vulnerabilities the model discovered and used two zero-day vulnerabilities. The system card section of 3 September 2026 says Astra is OpenAI's first model to reach the Critical level, and can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step.
First is OpenAI's own designation under its own framework, not a claim across labs. The 100% is the public ExploitBench result and the two zero-days came from the internal port, which are separate results. The page read does not state that the 100% was without safeguards, so that qualifier is unverified. All scores are vendor-reported and not independent. Delaying part of the release is supported by the post. The line about what models are stopped from doing is the author's opinion.
Watch next
- Independent evaluation of the cyber capability.
Sources
- Path to Astra (OpenAI, 1 Sep 2026)openai.com
- GPT-6 Astra system card, cybersecurity (OpenAI, 3 Sep 2026)deploymentsafety.openai.com
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 18 September 2026 at 23:48 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →