← Founder Notes
Archive

Openai designated gpt-6 astra as the first model at its critical cybersecurity threshold, meaning…

Yethikrishna ROriginal on Threads

openai designated gpt-6 astra as the first model at its critical cybersecurity threshold, meaning it can find unknown flaws and build exploits on well-protected systems without step by step guidance, and delayed part of the release to harden it. without safeguards it scored 100 percent on exploitbench and found two previously unknown vulnerabilities during evaluation.

the selling point of frontier models is no longer what they can do, it is what they are stopped from doing.

Context

OpenAI's post of 1 September 2026 says it believes Astra meets the Critical cybersecurity capability threshold under its Preparedness Framework, and that over several weeks it delayed parts of Astra's development and release while strengthening protections. It reports a perfect score of 100% on ExploitBench, and says that on an internal ExploitBench port of 20 recent V8 vulnerabilities the model discovered and used two zero-day vulnerabilities. The system card section of 3 September 2026 says Astra is OpenAI's first model to reach the Critical level, and can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step.

How it compares

First is OpenAI's own designation under its own framework, not a claim across labs. The 100% is the public ExploitBench result and the two zero-days came from the internal port, which are separate results. The page read does not state that the 100% was without safeguards, so that qualifier is unverified. All scores are vendor-reported and not independent. Delaying part of the release is supported by the post. The line about what models are stopped from doing is the author's opinion.

Watch next

  • Independent evaluation of the cyber capability.

Sources

  1. Path to Astra (OpenAI, 1 Sep 2026)openai.com
  2. GPT-6 Astra system card, cybersecurity (OpenAI, 3 Sep 2026)deploymentsafety.openai.com

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 18 September 2026 at 23:48 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/openai-designated-gpt-6-astra-as-the-first-DdcEAQ_kRwO" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Openai designated gpt-6 astra as the first model at its critical cybersecurity threshold, meaning…"></iframe>

More notes