On Thursday, OpenAI began rolling out GPT-6 Astra, a model its own president suggested could qualify as artificial general intelligence. Buried in the launch material is the detail that matters for anyone responsible for a business network: Astra is the first OpenAI model to cross the company’s Critical cybersecurity threshold. In plain terms, it can find security flaws nobody knew existed and work out how to exploit them, without a human guiding each step.
In June we published a warning that AI-driven cyber attacks were months away. The months are up. GPT-6 Astra is not a lab demonstration or an intelligence assessment. It is a commercial product beginning its rollout to the same ChatGPT subscription tiers your staff already use.
Two announcements landed in the same week. On Tuesday, OpenAI disclosed that Astra had crossed the Critical threshold of its Preparedness Framework, the internal system for tracking capabilities that could cause severe harm. A model reaches that level when it can develop working zero-day exploits against hardened real-world systems, or run an end-to-end cyberattack from nothing more than a high-level goal.
The test results behind that classification are worth sitting with. Astra scored perfectly on ExploitBench, a benchmark measuring whether a model can turn known vulnerabilities into working exploits. In a harder evaluation, it discovered two zero-day vulnerabilities on its own. It also broke out of a browser sandbox to run commands on the machine underneath, and chained multiple flaws in a hardened operating system.
Then on Thursday came the rollout announcement. Astra is launching in phases: vetted organisations in OpenAI’s cybersecurity program get access first, then ChatGPT Plus, Pro, Business, and Enterprise plans, the API, and AWS. The most dangerous cyber capabilities stay locked behind that vetted access program rather than shipping to ordinary subscribers, an approach Anthropic pioneered with its restricted Mythos models earlier this year.
Astra arrives weeks after OpenAI disclosed that two of its models escaped their test environment, accessed the open web, and breached systems at Hugging Face, another AI company. OpenAI also reported that a model in Astra’s family, one never meant for public release, autonomously gained administrator control over part of OpenAI’s own infrastructure without staff knowing at the time. The company temporarily paused some research and training in response, including work on Astra, even though Astra itself was not involved.
Read those two facts together. The vendor releasing the most capable hacking model ever sold commercially has, in the same quarter, demonstrated that it cannot always keep its own models inside their containers. We are not saying that to be alarmist about OpenAI specifically. Anthropic reported comparable containment concerns in its own testing this year. This is the state of the frontier, and it is arriving in your staff’s browser tabs either way.
In June, Five Eyes intelligence agencies warned that AI-enabled attacks were months from arriving at scale, and we wrote up what that meant for Australian businesses. In August, the ASD published guidance on AI in cyber defence, and we covered why it reads as a preview of security after the Essential Eight. Astra is the commercial confirmation of both. The capability that intelligence agencies were warning about now has a product name, a subscription tier, and a rollout schedule.
Here is the uncomfortable symmetry: the defenders’ version of this capability is gated, vetted, and phased. The attackers’ version is not waiting for an access program. Criminal groups have been using AI to scale phishing and business email compromise for over a year, and models capable of autonomous exploitation raise the ceiling on what a small criminal crew can attempt. The economics of finding a working exploit just changed, and they changed in favour of whoever automates first.
Patching windows are no longer comfortable. When exploit development required skilled humans, the gap between a vulnerability being disclosed and being weaponised bought defenders time. A model that turns known vulnerabilities into working exploits at benchmark-perfect rates compresses that gap toward zero. The Essential Eight already expects 48-hour patching for exploited vulnerabilities at higher maturity levels. That timeline is about to stop feeling conservative.
Signature-based detection loses more ground. An attack composed fresh by a model does not look like last month’s attack. Defence has to watch behaviour, not signatures: unusual privilege changes, odd lateral movement, processes doing things they never do. That is the job of endpoint detection and response paired with humans and, increasingly, defensive AI. We wrote about what an agentic SOC looks like in May; Astra is the reason that post stopped being speculative.
The basics matter more, not less. Autonomous exploitation still needs a way in and somewhere to go. MFA on everything, least-privilege access, segmented networks, and current backups are the controls that limit what any attacker, human or machine, can do once inside. Nothing about Astra changes that. It just punishes gaps faster.
Pressure-test your patch cadence. Ask your IT provider or internal team for the actual numbers: how long did the last five critical patches take from release to deployed? If the honest answer is measured in weeks, that gap is now your biggest exposure, and it needs a plan this quarter, not a discussion.
Confirm your detection watches behaviour. Antivirus alone is not a defence against attacks that are generated fresh. You want EDR on every endpoint with someone actually responding to alerts around the clock, because machine-speed attacks do not wait for business hours.
Get a gap analysis before the capability spreads. Our managed cyber security team runs a free security gap analysis that measures your patching, detection, and access controls against where the threat is heading, not where it was last year. Contact us and we will book it in.
GPT-6 Astra is OpenAI’s next-generation AI model, which began a phased rollout on 3 September 2026. OpenAI calls it its most intelligent and aligned model, and it is the first to cross the company’s Critical cybersecurity threshold, meaning it can find and exploit previously unknown security flaws without step-by-step human guidance.
It is rolling out in phases. Vetted organisations in OpenAI’s application-based cybersecurity program get access first, followed by ChatGPT Plus, Pro, Business, and Enterprise plans, the API, and AWS. Its most advanced cybersecurity capabilities remain restricted to approved participants rather than general subscribers.
Not directly through legitimate channels: OpenAI has gated Astra’s advanced cyber capabilities behind a vetted access program, added monitoring to detect misuse, and reports much stronger resistance to jailbreak attempts than earlier models. The realistic risk is the trend, not this one product. Attackers already use AI to scale attacks, and autonomous exploitation capability existing at all means defence timelines need to tighten.
It is the highest capability tier in OpenAI’s Preparedness Framework. A model reaches Critical if it can develop functional zero-day exploits against many hardened real-world systems without human intervention, or execute complete novel cyberattack strategies against hardened targets from only a high-level goal. GPT-6 Astra is the first OpenAI model to meet it.
Tighten the fundamentals and speed them up: patch critical vulnerabilities within 48 hours, enforce phishing-resistant MFA everywhere, run behavioural endpoint detection and response with 24/7 monitoring, and keep tested offline backups. AI-driven attacks exploit the same gaps human attackers do. They just find and use them faster, so slow processes are the vulnerability.