Back to Blog
Last updated Sep 08, 2026.

Three Cyber AI Models in 48 Hours, and None Is Generally Available

minutes read
Ali Ahmed

Ali Ahmed

AI Solutions Engineer, Cognilium AI

Share:
Three Cyber AI Models in 48 Hours, and None Is Generally Available
Anthropic, Google and OpenAI each shipped a cyber-capable frontier model inside 48 hours. All 3 sit behind a vetting gate, and none has a self-serve path.
cybersecurityAnthropicGoogleOpenAI

TL;DR

Between 1 and 2 September 2026, three labs announced their most capable security models: Claude Mythos 5.1, Gemini 3.8 Flash Cyber, and OpenAI Astra.

None of the three can be bought. Each sits behind a named vetting programme, and 0 of the 3 has a self-serve path.

OpenAI classed Astra at the Critical cybersecurity level under its Preparedness Framework, defined as a model that can "independently find and exploit zero-day vulnerabilities across many well-defended systems".

Google's Fairwind Program runs with more than 650 partners and three eligibility categories. Anthropic's Mythos 5.1 is "only available to a set of US organizations."

What you can buy this week is the softer half: Claude Fable 5.1 may now identify vulnerabilities defensively, with about 60% fewer false-positive safeguard interventions per session.

What actually happened this week?

Three announcements landed inside 48 hours, and read together they say something none of them says alone.

On 1 September, Anthropic published Claude Fable 5.1 and Claude Mythos 5.1. On 2 September, Google published Gemini 3.8 Flash and Gemini 3.8 Flash Cyber alongside the Fairwind Program, and OpenAI's Astra was reported at the Critical cybersecurity threshold under its own Preparedness Framework.

Each lab framed its release as a capability story. The more interesting story is the distribution decision underneath it, because all three made the same one and gave it three different names.

Who can actually get these models?

Nobody, on a credit card. Here is the gate on each.

Lab. Model. Named gate. Who qualifies. Self-serve

Anthropic. Claude Mythos 5.1. Cyber Verification Program (CVP), Life Sciences Verification Program (LSVP). Vetted cyberdefenders and life scientists, "only available to a set of US organizations". No

Google. Gemini 3.8 Flash Cyber. Fairwind Program. Governments and national cyber authorities, critical infrastructure operators, core technology platforms. No

OpenAI. Astra, full cyber capability. Daybreak Blue, then the wider Daybreak programme. Testers first, "will not be widely available at launch". No

Three labs, three gates, zero of them open. That is a coordinated posture arrived at independently, which is usually a sign the reasoning is the same in all three rooms.

Google is the most explicit about the terms. Fairwind participants must accept "strict operational standards, including limiting access to employees within their internal cybersecurity, incident response, or penetration testing teams and deploying protections like multi-factor authentication." That is not a licence agreement. It is closer to a clearance.

What does "Critical" actually mean?

It means the model can do the attack on its own.

OpenAI defines the Critical cybersecurity capability level as a model that can "independently find and exploit zero-day vulnerabilities across many well-defended systems, or carry out a complete cyberattack against a hardened target from only a high-level instruction." Astra is the first OpenAI model placed in that category.

The evidence behind the classification is specific rather than rhetorical. Astra scored a perfect result on ExploitBench, which measures turning known vulnerabilities into working exploits. In a separate evaluation over more recently disclosed flaws it found 2 zero-day vulnerabilities on its own. It escaped a browser sandbox to run commands on the host, and chained several operating-system flaws to reach root.

The jailbreak number everyone is reading backwards

OpenAI reports that Astra declines 91.5% of cyber-related jailbreak attempts, against 59% for GPT-5.6 Sol. Most coverage reads that as a 32.5 point improvement, which is true and close to useless.

The number a defender cares about is what gets through.

GPT-5.6 Sol 100% - 59.0% = 41.0% of attempts succeed

Astra 100% - 91.5% = 8.5% of attempts succeed

5 / 41.0 = 0.207

So the share of jailbreak attempts that land falls to about 21% of what it was, a 79% reduction. Per 1,000 attempts against the model, 410 used to get through and 85 now do. That is 325 fewer successful jailbreaks per thousand tries.

The same arithmetic sets the ceiling honestly. 8.5% is not zero. At an attacker volume of 1,000 attempts a day, a model at Critical capability still yields on 85 of them.

Is the Google model actually better, or just gated?

The published comparisons are unusually concrete for a launch post, so they are worth doing the arithmetic on.

Google reports Gemini 3.8 Flash Cyber producing 2.6 times more correct patches to Chrome vulnerabilities than the best commercial models it tested. Against Wiz's workload it reports 7.5% to 9.7% higher recall at 2.3 to 5.2 times lower cost.

Those two Wiz figures are usually quoted separately. Combined, they give the number a security budget actually runs on, which is cost per vulnerability found:

relative cost 1 / 2.3 = 0.435 … 1 / 5.2 = 0.192

recall multiplier 1.075 … 1.097

cost per finding 0.435 / 1.075 = 0.405 … 0.192 / 1.097 = 0.175

So a finding costs between 17.5% and 40.5% of what it did, which is roughly 60% to 82% cheaper per vulnerability actually found. On internal vulnerability discovery across 20 programming languages Google reports a success rate above 70%, and 47.2% pass@1 on CWE-Bench patching.

Those are strong numbers. They are also numbers you cannot act on unless you are a government, a critical infrastructure operator or a core technology platform.

What can an ordinary security team buy today?

The generally available half, which is not nothing.

Available now. What changed

Claude Fable 5.1. May now "be used for identifying software vulnerabilities" for defensive work. Cyber safeguards block 60% fewer false positives, about 60% fewer interventions per session. Costs fall about 25% on typical workloads and up to about 45% on complex coding. Cache reads drop 75%, to $0.25 per million tokens

Gemini 3.8 Flash. $0.75 per million input tokens and $3.75 per million output, an introductory rate through 31 December 2026. 54.9% on HLE-Verified

Note the second row carefully, because it has a clock on it. The introductory rate rises to $1.50 and $7.50 afterwards, which is exactly double. As of today there are 114 days of introductory pricing left.

A workload of 100 million input and 20 million output tokens a month costs $150 today and $300 from January, on identical usage. Any 2027 budget built on this month's rate card is out by a factor of 2.

Why does the gating matter to a business that will never qualify?

Because the pattern is arriving in business software, and it changes what "our vendor has the best model" means.

Every one of these three labs supplies models into enterprise platforms. When the most capable version of a model is reachable only through a vetting programme, the version inside your ERP, your service desk or your code assistant is by definition not that one. The marketing sits at the frontier and the deployment sits behind the gate, and both statements are true at once.

We looked at how Anthropic built Enterprise Frontier Safeguards with its own customers last week, and at the five-way frontier race before that. This week completes the shape. The competition is still on capability. The release is increasingly not.

There is a second-order effect worth naming. A vetting gate is a list, and a list is a governance artefact. Google's more than 650 Fairwind partners are, collectively, a register of who a major lab considers a trusted defender. No regulator asked for that register. It exists because three companies each decided, separately, that the alternative was worse.

What should a security or platform team do about it?

Check whether you qualify before assuming you do not. The Fairwind categories reach further than they first read. "Core technology platforms" is a broad phrase, and Anthropic's CVP is open to application rather than invitation only.

Ask your AI vendor which model version you are actually served. For a model family with gated and ungated variants, the contract rarely says. It is a fair question and the answer is a fact, not an opinion.

Re-baseline your token budget before January. One of the two generally available models on this list doubles in price in 114 days. That is a calendar item, not a negotiation.

Read the refusal rate as a leak rate. A model that declines 91.5% of attempts still yields on 8.5%. If your control assumes the model is the control, the arithmetic above is your exposure.

FAQ

Can I buy access to Gemini 3.8 Flash Cyber?

No. It is limited to trusted defenders through the Fairwind Program, which is applied for rather than purchased.

Is Claude Mythos 5.1 available outside the United States?

Not currently. Anthropic states it is "only available to a set of US organizations."

When will Astra's full cybersecurity features be generally available?

No date has been given. OpenAI said full cybersecurity capabilities "will not be widely available at launch", with early access through Daybreak Blue and wider availability through Daybreak afterwards.

Does Claude Fable 5.1 write exploits?

No. It may identify software vulnerabilities for defensive work. Penetration testing, exploit generation and binary-based vulnerability scanning remain outside what it will do.

What is the difference between Fable 5.1 and Mythos 5.1?

Anthropic describes them as the same model with different levels of safeguards. Fable is generally available. Mythos is reachable only through the verification programmes.

The last mile

The interesting fact about this week is not that three models got better at security. It is that three companies looked at the same capability and independently decided that selling it was the wrong move.

That decision has a cost, and it is being paid by the defenders who do not make the list. The strongest tool for finding a flaw in your own code is now something you apply for. For most organisations the practical answer is not access to a frontier cyber model. It is knowing exactly which decisions inside your own systems are being made by a model, on what evidence, and whether you could explain any one of them to an auditor. That layer, sitting inside the systems a business already runs on, is where Cognilium works, and it does not require anyone's clearance.

Share this article

Share:

Weekly AI engineering brief

One email a week. New model releases, agent patterns, and lessons from production systems we ship.

No spam, no client data sales. Unsubscribe any time.

Ali Ahmed

Ali Ahmed

AI Solutions Engineer, Cognilium AI

Ali Ahmed is an AI Solutions Engineer at Cognilium AI.