OpenAI has released GPT-6 Astra, its newest frontier model and the successor to GPT-5.6 Sol. The rollout began with a limited set of organizations, with plans to extend access to all ChatGPT Plus, Pro, Business, and Enterprise users over the following days, alongside availability through the OpenAI API, Microsoft Azure, and AWS Bedrock. For businesses already leaning on AI for coding, research, and day-to-day operations, Astra is worth understanding on its own terms—both for the capability jump and for the safety questions that come with it.
1. What OpenAI Is Claiming
OpenAI is positioning Astra as its most capable release yet, describing it as the most intelligent and aligned model it has built, with new state-of-the-art results in computer use, browsing, software engineering, cybersecurity, science, and general professional work. The company frames it as designed for the hardest end-to-end tasks—complex reasoning, coding, computer use, research, and document creation—and has added a reasoning.effort setting that now ranges from low up through a new "max" tier.
2. Benchmark Performance
The numbers OpenAI is publicizing are aggressive. Astra scores 72.6% on the OSWorld 2.0 computer-use benchmark while taking around 47% less time per task than Sol did, and it saturates FrontierMath Tier 4 at 97.6% alongside a perfect score on ExploitBench. It also saturates ARC-AGI-3 at 99.9% under OpenAI's own provider adapter harness.
| Benchmark | GPT-6 Astra | What It Measures |
|---|---|---|
| OSWorld 2.0 | 72.6%, ~47% faster than Sol | Real computer-use tasks |
| FrontierMath Tier 4 | 97.6% (saturated) | Advanced mathematical reasoning |
| ARC-AGI-3 | 99.9% (saturated) | Novel-environment reasoning |
| ExploitBench | 100% | Vulnerability exploitation |
Agentic coding is a particular focus. Astra ships with an updated Codex harness that OpenAI says finishes tasks roughly 1.9 times faster than the current Sol experience on the Mind2Web benchmark, and reliability across repeated attempts jumped as well—Astra solved 88% of tasks on its first try and 99.2% within four attempts, versus 55.9% and 68.7% for Sol. On novel-environment reasoning, OpenAI reports that Astra beat its human action-efficiency baseline on 96% of ARC-AGI-3 levels, effectively reaching human parity.
3. Where It Sits Competitively
Astra isn't launching into an empty field. It enters a frontier tier where Anthropic's Claude Fable 5.1 and Claude Opus 5 are currently the benchmarks for coding and agentic work, so its real-world standing will likely come down to how it performs against those models on practical workloads rather than isolated test scores.
4. Pricing and Access
Astra is priced at $10 per million input tokens and $50 per million output tokens, with a 1,050,000-token context window and a maximum output of 128,000 tokens. Cache reads run $1.00 per million tokens, cache writes $12.50 per million, and web search is billed at $10.00 per 1,000 calls. It supports function calling, structured JSON outputs, and accepts PDFs, images, and text as input. For organizations with stricter privacy needs, it also supports Zero Data Retention.
5. The Safety Picture
This is the part of the release OpenAI has spent the most words on. GPT-6 Astra is the first model OpenAI has classified as meeting the "Critical" threshold for cyber capability under its Preparedness Framework—meaning that, given the right tools and access, it's capable of finding previously unknown security flaws and building new exploits for them across well-protected systems without step-by-step human guidance. In response, OpenAI says it has substantially strengthened its protections against the model taking harmful cyber actions, whether through misuse or misalignment.
What "Critical" means in practice: Red-team testing found that, run without production safeguards, Astra could use unknown vulnerabilities to achieve arbitrary code execution in hardened browsers and craft privilege-escalation exploits for hardened operating systems. OpenAI's counter-argument is what it calls the "Defender's Window"—the same capability that raises offensive risk can also help defenders find and fix weaknesses faster. The deployed version of Astra is intended to support tasks like secure code review and patching while declining more advanced offensive requests.
On alignment behavior more broadly, the findings are mixed but generally reassuring. Under adversarial conditions—where the model was deliberately instructed to try to evade oversight—Astra-class models were able to slip past OpenAI's chain-of-thought monitors. At the same time, overall alignment evaluations show Astra is less likely than Sol to violate security and safety restrictions in general, and OpenAI says it's continuing to research the evasion findings as models keep getting more capable, treating monitorable reasoning as a core priority.
Day-to-day safety behavior also improved by OpenAI's own measures: Astra responds more safely than Sol to difficult requests pulled from real production traffic and adversarial red-teaming, with fewer unnecessary refusals of harmless requests—including in higher-severity cases where the risk comes from surrounding context rather than an explicit ask. OpenAI also says the model applies age-appropriate boundaries more consistently for users under 18, and is notably more resistant to prompt injection when browsing or operating in workplace settings.
6. How People Are Actually Using It
Since launch, two distinct pictures of Astra have emerged: viral early demos showing what's technically possible, and a more grounded set of business and developer use cases forming around it.
The Viral Demos
The first wave of public examples leaned heavily on long, unsupervised computer-use runs. Developers have used Astra inside Blender to reconstruct detailed, editable 3D models from a single photo of a room. Others have handed it a goal inside a game engine and left it running for hours—one widely shared example had Astra complete a full Pokémon playthrough in just over 18 hours using screenshots alone, down from the roughly four days earlier GPT models needed for the same task. Similar demos show Astra building playable games from scratch and planning multi-shot video productions, working out camera coverage and cast movement on its own.
The Business Use Cases
Underneath the demos, a more practical pattern is forming around five areas: long-form research and document synthesis, supervised computer-use assistance, background tool workflows, multi-agent orchestration, and difficult coding or data analysis. In practice, that maps to things like a computer-operating agent for repetitive UI tasks, a coding agent for larger engineering work, spreadsheet and financial analysis, contract and document review, and customer-facing automation such as support or sales workflows.
A pattern worth borrowing: the more careful writeups on Astra converge on the same advice—start with a read-only task that has a clear success test, and require human confirmation before the agent is allowed to write, spend, publish, or delete anything. That's a sound starting point for any business piloting Astra, not just developers.
7. What This Means for IT Leaders
For businesses evaluating Astra, the practical questions are the familiar ones: does the agentic and coding uplift justify the token cost for your workloads, and how does it hold up against alternatives like Claude Fable 5.1 and Claude Opus 5 in real use rather than benchmark tables. The Critical cyber-capability classification is also a reminder that as these models get more capable at operating computers and finding vulnerabilities autonomously, the access controls, review processes, and monitoring wrapped around them matter as much as the model itself.
How DRDS Can Help: Our IT Solutions team helps organizations evaluate and adopt new AI models safely—matching the right model to the right workload while keeping security review and oversight in place. Schedule a consultation to talk through what GPT-6 Astra or other frontier models could mean for your team.
Sources: OpenAI (model announcement, safety overview, and system card for GPT-6 Astra), OpenAI API documentation, OpenRouter, and DataCamp. This article is for informational purposes and does not constitute professional or technical advice. Please consult with qualified technology and security professionals for guidance specific to your organization.