GPT-6 Astra’s “Critical” Cyber Risk Rating Raises New Security Concerns
OpenAI's newest ChatGPT model, GPT-6 Astra, is the first model the company has ever rated "Critical" for cybersecurity risk — but that rating describes what the model could theoretically do without safeguards, not what happens when you open ChatGPT and type a normal question. If you use ChatGPT for writing, research, or everyday tasks, nothing about your experience changes because of this label.
Key Takeaway: "Critical" is a formal safety classification under OpenAI's own Preparedness Framework, not a warning that Astra is dangerous to use casually. It means the underlying model, if stripped of its guardrails, is capable of finding and exploiting unknown security flaws on its own. The version you actually get in ChatGPT ships with extra restrictions specifically because of this rating.
If you've seen the "$0 to critical" headline circulating this week, it's accurate — and it's also doing a lot of work to make Astra sound scarier to use than it actually is. Here's what the rating actually breaks down to, and why OpenAI shipped it anyway.
What Actually Happened
OpenAI began rolling out GPT-6 Astra in the first week of September 2026, calling it the most capable model it has ever broadly deployed. Alongside the launch, OpenAI published a safety overview stating that Astra is the first model to cross the "Critical" threshold for cybersecurity capability under its internal Preparedness Framework — the system OpenAI uses to grade how risky a model's raw abilities are before deciding how to release it.
That framework has four tiers: low, medium, high, and critical. OpenAI's previous flagship, GPT-5.6 Sol, sat at "high." We covered that release in detail in our GPT-5.6 Explained: Sol, Luna, Terra & Pricing guide, including the government review process Sol went through before public release. Astra is the first model to go one step further, into "critical" territory.
What "Critical" Actually Means
According to OpenAI's own published safety documentation, a model at the critical tier can, with the right tools and access, find previously unknown security flaws and build working exploits for them across well-protected systems — without a person walking it through each step. In testing without any production safeguards attached, Astra reportedly scored 100% on ExploitBench (up from 78.5% for GPT-5.6 Sol) and found two real, previously unknown vulnerabilities during evaluation, which OpenAI says it has since disclosed to the affected software makers.
That's a genuinely different level of capability than earlier models. But it's a description of what Astra can do in a stripped-down testing environment — not a description of the ChatGPT app sitting on your phone right now.
What Changes for a Regular ChatGPT User — and What Doesn't
This is the part most coverage of Astra has skipped, because most of it was written for security teams, not for the millions of people who just use ChatGPT to draft emails or research topics. Here's the honest breakdown:
What doesn't change: Your day-to-day ChatGPT experience — writing, summarizing, coding help, general questions — works the same as it did on the previous model. The shipped consumer version of Astra still refuses to help with exploit development, malware, or attacks on systems you don't own, the same as every OpenAI model before it. OpenAI says Astra actually performs better on alignment and safety testing than its predecessor, not worse.
What does change: Access to Astra's more advanced capabilities is more tightly controlled than any previous release. Enterprise accounts have to manually turn Astra on — it's off by default. Anyone doing legitimate security work, like penetration testing or vulnerability research, needs to apply for OpenAI's separate "Daybreak" program to get a less restricted version, rather than being able to do that through the normal consumer product.
| Preparedness Framework Tier | In Plain English | Example Model |
|---|---|---|
| Low | No meaningful uplift for cyberattacks beyond normal tools | Older GPT-4-class models |
| Medium | Can meaningfully speed up a skilled person's work | Earlier GPT-5 models |
| High | Can help less-skilled people do advanced attacks with guidance | GPT-5.6 Sol |
| Critical | Can find and exploit unknown flaws largely on its own | GPT-6 Astra |
Why OpenAI Shipped It Anyway
OpenAI's argument, laid out in its own safety materials, is that the same capability that makes Astra risky in the wrong hands also makes it valuable in the right ones — specifically for the defenders trying to find and patch flaws before attackers do. That's the reasoning behind the Daybreak program, which OpenAI has paired with a separate billion-dollar commitment to subsidize access for smaller security teams that couldn't otherwise afford frontier-model tools. In effect, OpenAI is betting that giving defenders a head start with this capability outweighs the risk of it eventually leaking into the wrong hands through jailbreaks or misuse.
Whether that bet pays off is genuinely unresolved, and independent researchers will be watching for real-world jailbreak attempts against Astra in the coming months, the same way they've tested every major model release before it.
How Astra Compares to OpenAI's Last Flagship
If you're deciding whether to pay attention to this release at all, it helps to see it next to the model it replaces. We broke down how GPT-5.6 Sol stacked up against a rival flagship in our Grok 4.6 vs GPT-5.6 comparison — Astra is a meaningful step beyond that generation specifically in computer-use and cybersecurity-adjacent tasks, rather than a broad, across-the-board leap in every category.
Pro Tip: If you run a small business or manage IT for a team, this is a good moment to check whether your organization has Astra enabled by default in your workspace admin settings — remember, OpenAI ships it off by default for enterprise accounts, so someone has to turn it on deliberately.
The Bottom Line
GPT-6 Astra crossing OpenAI's "Critical" cybersecurity threshold is a real, notable safety milestone — it's the clearest public signal yet that AI models are approaching genuinely dangerous offensive-security capability. But for the overwhelming majority of ChatGPT users, this is a story about what OpenAI is doing behind the scenes to manage that capability, not a change to what you'll experience the next time you open the app.
Sourcing note: This article draws on OpenAI's own published safety overview and system card for GPT-6 Astra, along with independent reporting from CSO Online on the model's launch details and benchmark figures. Some specifics of OpenAI's internal testing methodology have not been independently verified beyond what the company has disclosed publicly.
