GPT-6 Astra’s “Critical” Cyber Risk Rating Raises New Security Concerns

GPT-6 Astra critical cyber risk

OpenAI's newest ChatGPT model, GPT-6 Astra, is the first model the company has ever rated "Critical" for cybersecurity risk — but that rating describes what the model could theoretically do without safeguards, not what happens when you open ChatGPT and type a normal question. If you use ChatGPT for writing, research, or everyday tasks, nothing about your experience changes because of this label.

Key Takeaway: "Critical" is a formal safety classification under OpenAI's own Preparedness Framework, not a warning that Astra is dangerous to use casually. It means the underlying model, if stripped of its guardrails, is capable of finding and exploiting unknown security flaws on its own. The version you actually get in ChatGPT ships with extra restrictions specifically because of this rating.

If you've seen the "$0 to critical" headline circulating this week, it's accurate — and it's also doing a lot of work to make Astra sound scarier to use than it actually is. Here's what the rating actually breaks down to, and why OpenAI shipped it anyway.

What Actually Happened

OpenAI began rolling out GPT-6 Astra in the first week of September 2026, calling it the most capable model it has ever broadly deployed. Alongside the launch, OpenAI published a safety overview stating that Astra is the first model to cross the "Critical" threshold for cybersecurity capability under its internal Preparedness Framework — the system OpenAI uses to grade how risky a model's raw abilities are before deciding how to release it.

That framework has four tiers: low, medium, high, and critical. OpenAI's previous flagship, GPT-5.6 Sol, sat at "high." We covered that release in detail in our GPT-5.6 Explained: Sol, Luna, Terra & Pricing guide, including the government review process Sol went through before public release. Astra is the first model to go one step further, into "critical" territory.

What "Critical" Actually Means

According to OpenAI's own published safety documentation, a model at the critical tier can, with the right tools and access, find previously unknown security flaws and build working exploits for them across well-protected systems — without a person walking it through each step. In testing without any production safeguards attached, Astra reportedly scored 100% on ExploitBench (up from 78.5% for GPT-5.6 Sol) and found two real, previously unknown vulnerabilities during evaluation, which OpenAI says it has since disclosed to the affected software makers.

That's a genuinely different level of capability than earlier models. But it's a description of what Astra can do in a stripped-down testing environment — not a description of the ChatGPT app sitting on your phone right now.

What Changes for a Regular ChatGPT User — and What Doesn't

This is the part most coverage of Astra has skipped, because most of it was written for security teams, not for the millions of people who just use ChatGPT to draft emails or research topics. Here's the honest breakdown:

What doesn't change: Your day-to-day ChatGPT experience — writing, summarizing, coding help, general questions — works the same as it did on the previous model. The shipped consumer version of Astra still refuses to help with exploit development, malware, or attacks on systems you don't own, the same as every OpenAI model before it. OpenAI says Astra actually performs better on alignment and safety testing than its predecessor, not worse.

What does change: Access to Astra's more advanced capabilities is more tightly controlled than any previous release. Enterprise accounts have to manually turn Astra on — it's off by default. Anyone doing legitimate security work, like penetration testing or vulnerability research, needs to apply for OpenAI's separate "Daybreak" program to get a less restricted version, rather than being able to do that through the normal consumer product.

Preparedness Framework Tier In Plain English Example Model
Low No meaningful uplift for cyberattacks beyond normal tools Older GPT-4-class models
Medium Can meaningfully speed up a skilled person's work Earlier GPT-5 models
High Can help less-skilled people do advanced attacks with guidance GPT-5.6 Sol
Critical Can find and exploit unknown flaws largely on its own GPT-6 Astra

Why OpenAI Shipped It Anyway

OpenAI's argument, laid out in its own safety materials, is that the same capability that makes Astra risky in the wrong hands also makes it valuable in the right ones — specifically for the defenders trying to find and patch flaws before attackers do. That's the reasoning behind the Daybreak program, which OpenAI has paired with a separate billion-dollar commitment to subsidize access for smaller security teams that couldn't otherwise afford frontier-model tools. In effect, OpenAI is betting that giving defenders a head start with this capability outweighs the risk of it eventually leaking into the wrong hands through jailbreaks or misuse.

Whether that bet pays off is genuinely unresolved, and independent researchers will be watching for real-world jailbreak attempts against Astra in the coming months, the same way they've tested every major model release before it.

Warning: The "Critical" label is not a reason to stop using ChatGPT normally, and it doesn't mean your account or device is at higher risk simply because Astra exists. The real risk this classification is meant to manage is misuse by sophisticated bad actors attempting to bypass safeguards at scale — not casual, everyday usage.

How Astra Compares to OpenAI's Last Flagship

If you're deciding whether to pay attention to this release at all, it helps to see it next to the model it replaces. We broke down how GPT-5.6 Sol stacked up against a rival flagship in our Grok 4.6 vs GPT-5.6 comparison — Astra is a meaningful step beyond that generation specifically in computer-use and cybersecurity-adjacent tasks, rather than a broad, across-the-board leap in every category.

Quick Win: Curious which model you're actually using right now? Open ChatGPT, tap your model name at the top of the chat window, and check the dropdown. If Astra hasn't reached your account yet, your existing model will keep working exactly as before — there's no action required on your end.
Pro Tip: If you run a small business or manage IT for a team, this is a good moment to check whether your organization has Astra enabled by default in your workspace admin settings — remember, OpenAI ships it off by default for enterprise accounts, so someone has to turn it on deliberately.

The Bottom Line

GPT-6 Astra crossing OpenAI's "Critical" cybersecurity threshold is a real, notable safety milestone — it's the clearest public signal yet that AI models are approaching genuinely dangerous offensive-security capability. But for the overwhelming majority of ChatGPT users, this is a story about what OpenAI is doing behind the scenes to manage that capability, not a change to what you'll experience the next time you open the app.

Sourcing note: This article draws on OpenAI's own published safety overview and system card for GPT-6 Astra, along with independent reporting from CSO Online on the model's launch details and benchmark figures. Some specifics of OpenAI's internal testing methodology have not been independently verified beyond what the company has disclosed publicly.


Author Image

Hardeep Singh

Hardeep Singh is a tech and money-blogging enthusiast, sharing guides on earning apps, affiliate programs, online business tips, AI tools, SEO, and blogging tutorials. About Author.

Previous Post