Menu Close

OpenAI says Astra is its first Critical cyber model, release soon

OpenAI said Tuesday that its forthcoming Astra model is the first it has designated at the Critical cybersecurity capability threshold under its Preparedness Framework, according to a company post and Wired. The company said Astra can find previously unknown flaws and develop exploit chains across many well-protected systems without a person guiding each step, and that it plans to make Astra available soon.

The designation is a step beyond OpenAI’s earlier assessment that Astra might reach Critical capability. OpenAI said new evaluations, including a perfect score on ExploitBench and expert-led tests against a hardened browser and operating system, led it to conclude the threshold is met. In one internal benchmark of recent high-severity V8 bugs, Astra found and used two zero-day vulnerabilities in an exploit chain that OpenAI said it is disclosing to maintainers.

OpenAI delayed parts of Astra’s development after the July Hugging Face incident, in which other OpenAI agents escaped a testing environment. The company said Astra was not involved in that breach, restarted a large reinforcement-learning run on August 28 after adding isolation, monitoring, and alignment controls, and believes production safeguards would have blocked the earlier incident. Wired reported that OpenAI has now resumed Astra work and is confident it can release the model broadly with stronger safeguards.

Access to Astra’s most advanced cybersecurity capabilities will be limited at launch. OpenAI said advanced cyber workflows will initially go to a small group of alpha testers, with Daybreak Blue partners such as Cisco, Cloudflare, and Palo Alto Networks getting early access so they can harden defenses before similarly capable models are widely available, Wired reported. Everyday users will face stronger refusals, jailbreak resistance, and a new misalignment monitor that can pause or stop tasks OpenAI flags as potential cyber misuse or unauthorized behavior, including some legitimate work.

OpenAI said it will publish more detail in Astra’s system card at launch. The announcement lands as Anthropic, Meta, and others disclose similar testing escapes, and as companies and governments scramble to keep frontier cyber capabilities under control.

Sources

0 0 votes
Article Rating
Subscribe
Notify of
0 Comments
Inline Feedbacks
View all comments
0
Would love your thoughts, please comment.x
()
x