On September 1, OpenAI confirmed that Astra has reached its "Critical" cybersecurity threshold—the first model ever to do so. It will launch soon, but its most advanced capabilities will be restricted to a small group of testers. The company says Astra can find and exploit zero-day vulnerabilities without human guidance. It is also the most "aligned" model OpenAI has built. And the most dangerous.
OpenAI's Preparedness Framework defines "Critical" cybersecurity capability as a model that can autonomously identify and develop functional zero-day exploits in hardened real-world systems without human intervention, or execute novel end-to-end cyberattacks given only a high-level strategic goal. Previous models, including GPT-5.6-Sol, were assessed at the "High" threshold—not "Critical." Astra is the first.
Astra Can Find Zero-Days. No Human Guidance Needed.
In testing, Astra outperformed GPT-5.6 Sol on ExploitBench, a benchmark containing 20 high-severity vulnerabilities. It discovered and used two zero-day vulnerabilities as part of an exploit chain—OpenAI is now disclosing them to the maintainers.
The model's capabilities extend beyond benchmarks. OpenAI research scientist Fouad Matin captured the tension: "We believe these capabilities can help defenders discover and fix serious weaknesses, but without proper safeguards, they could also make attackers more efficient—which is exactly what we're trying to prevent."
Astra is also more likely to refuse harmful requests than GPT-5.6 Sol. In one evaluation, Astra refused 91.5% of malicious requests compared to 59% for GPT-5.6 Sol. But there's a trade-off: Astra might be too cautious and refuse legitimate cybersecurity work, like helping patch a vulnerability.
OpenAI Is Launching It. But Only for a Select Few.
Astra's release has already been delayed by "a certain number of weeks because everything was paused after Hugging Face, and then we took extra time to make sure that what we're launching is safe," a company spokesperson said.
When Astra launches, its most advanced cybersecurity capabilities will be available only to a small group of "alpha testers"—including U.S. government agencies and companies in OpenAI's trusted access program for cybersecurity. Later, access will expand through the "Daybreak Blue" program for defensive cybersecurity work. OpenAI acknowledged this approach may slow legitimate defensive work.
The company also said it will not release the model's cybersecurity evaluation code or benchmarks, citing the risk that doing so could help attackers.
Astra Wasn't the Breach Model. But It Learned From It.
Astra was not involved in the July Hugging Face breach. But OpenAI says it has incorporated lessons from that incident into Astra's safeguards, including training the model to more reliably refuse harmful cyber requests, adding protections against misuse, and building monitoring that can stop unauthorized activity.
The company has also tightened secure sandboxes and introduced monitoring that can identify and stop potentially unauthorized behavior. OpenAI's Preparedness Framework was published in late 2023 as an internal guide for identifying frontier model risks and deciding what safeguards to apply. The Astra designation is the first time those safeguards have been triggered at the highest level.

OpenAI is launching a model it believes could autonomously hack real systems. It is restricting that capability to a small group of testers and hoping the safeguards hold. The company is open about the contradiction. As its blog post put it: "We hope you enjoy our new model, and we hope the world continues to take what's happening in AI extremely seriously."
P.S. The quiet detail in the announcement: OpenAI says it believes "the world will need aligned AI to manage the future phases of this transition." Astra is the first model that is both capable enough to be dangerous and aligned enough to—hopefully—be trusted. That trust is the whole point. It's also the biggest unanswered question.
Frequently Asked Questions
Q: What is OpenAI's Astra model?
A: Astra is OpenAI's upcoming flagship AI model. On September 1, 2026, OpenAI announced it is the first model to reach its "Critical" cybersecurity threshold under the company's Preparedness Framework . This means Astra can autonomously identify and exploit zero-day vulnerabilities in hardened systems without human guidance .
Q: What does "Critical" cybersecurity threshold mean?
A: Under OpenAI's Preparedness Framework, "Critical" means the model can identify and develop functional zero-day exploits across multiple hardened real-world systems without human intervention, or execute novel end-to-end cyberattacks given only a high-level goal . Previous models, including GPT-5.6-Sol, were rated "High"—not "Critical." Astra is the first .
Q: When will Astra be released?
A: OpenAI says Astra will be available "soon" . The release has already been delayed by "a certain number of weeks" following the Hugging Face incident to strengthen safety measures . The most advanced cybersecurity capabilities will initially be limited to a small group of testers .
Q: What can Astra actually do?
A: In internal evaluations, Astra scored 100% on ExploitBench, a test of LLM ability to hack known vulnerabilities . It discovered and chained two zero-day vulnerabilities in an internal test and is disclosing them to the maintainers . It can find and exploit security flaws without a person guiding each step .
Q: Who gets access to Astra's most advanced capabilities?
A: Full access will be limited to a small group of "alpha testers," including U.S. government agencies and companies in OpenAI's trusted access program for cybersecurity . Access will expand later through the "Daybreak Blue" program for defensive cybersecurity work .
Q: Was Astra involved in the Hugging Face breach?
A: No. OpenAI confirmed Astra was not involved in the July Hugging Face incident . However, the company incorporated lessons from that incident into Astra's safeguards, including training the model to refuse harmful cyber requests and adding monitoring to stop unauthorized activity .
Q: Is Astra safer than previous models?
A: OpenAI says Astra is its "most aligned model to date." In one evaluation, Astra refused 91.5% of malicious requests compared to 59% for GPT-5.6 Sol . However, this creates a trade-off: Astra might be too cautious and refuse legitimate cybersecurity work .
Q: What is the ExploitBench score?
A: Astra scored 100% on the public ExploitBench benchmark. In an internal version with 20 high-severity V8 vulnerabilities, it achieved roughly 39% arbitrary code-execution rate and discovered two zero-day vulnerabilities .
Q: Why is OpenAI limiting access to Astra's capabilities?
A: OpenAI says Astra's capabilities could help defenders find vulnerabilities, but without safeguards could also make attackers more efficient . The company is balancing defensive benefits against misuse risk . OpenAI researcher Fouad Matin said this is "exactly what we're trying to prevent" .
Q: What is the significance of this announcement?
A: Astra is the first model OpenAI has ever designated as "Critical" risk—a watershed moment in AI development . It raises new questions about how to safely deploy models that can autonomously hack systems . The company's approach—launching but restricting capabilities—sets a precedent for future frontier model releases.
