
OpenAI has introduced GPT-6 Astra, a new frontier model the company calls its most intelligent and aligned system to date. Astra succeeds GPT-5.6 Sol as OpenAI’s flagship, and it lands into an already crowded field of frontier models, including Anthropic’s Claude Fable 5.1 and Claude Opus 5, and Google’s Gemini 3.8 Flash.
Astra’s headline claims are bold: it saturates several of the industry’s hardest benchmarks, sets new records in computer use and coding, and is the first OpenAI model to cross the Critical threshold for cybersecurity capability under the company’s own Preparedness Framework. This guide breaks down what GPT-6 Astra actually is, what is new about it, how it performs against the competition, what it costs, and where its claims come with fine print worth knowing before you build on it.
What Is GPT-6 Astra?
GPT-6 Astra is OpenAI’s new frontier flagship model, released on September 5, 2026. It is designed around three core areas: state of the art computer and browser use, a step change in professional knowledge work such as documents, spreadsheets, and presentations, and a major jump in cybersecurity capability, one large enough that it crosses the Critical threshold defined in OpenAI’s Preparedness Framework.
Astra is rolling out in stages, first to a limited set of organizations and then, over the following days, to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API, Microsoft Azure, and AWS Bedrock. Pro, Business, and Enterprise users also get access to a separate GPT-6 Astra Pro tier, which runs the same underlying model with its reasoning mode set for higher quality responses on complex tasks. Enterprise administrators can turn Astra on for their workspace, though it is off by default at launch. The model supports Zero Data Retention for eligible API customers, and OpenAI says it is also testing a Private Safety Processing option intended to strengthen safety monitoring while preserving customer privacy. In the API, developers can call it using the model name gpt-6-astra.
What Is New With GPT-6 Astra
Best in class computer use
GPT-6 Astra marks a new frontier in the speed, accuracy, and safety of computer use, according to OpenAI. It can handle tedious, multistep desktop work such as filling out online forms, updating records in a CRM, organizing a calendar, conducting online research, drafting summaries, analyzing scientific data, generating plots, building and testing a website, and troubleshooting problems on screen.
On OSWorld 2.0, a benchmark that measures whether an agent can complete real desktop tasks like navigating apps and manipulating files, Astra scores 72.6 percent while taking roughly 40 minutes per task, compared with GPT-5.6 Sol’s 65.7 percent at roughly 75 minutes, an improvement description independent commentators have summed up as about 47 percent less time for a meaningfully higher score. On ScreenSpot Pro, which tests how well a model can locate UI elements on a screen without using tools, Astra hits 92.7 percent against Sol’s 76.9 percent. Alongside Astra, OpenAI also updated its Codex harness, which it says delivers 1.9 times faster task completion than the current GPT-5.6 Sol experience on the Mind2Web benchmark.
A step change in professional work
GPT-6 Astra is trained to produce polished, business ready documents, spreadsheets, and presentations that follow existing templates rather than generic first drafts, matching a user’s writing and visual style while pulling in only the context that actually matters to the task. With Sites in ChatGPT, Astra can create, host, and share websites, web apps, and games directly from a prompt, and OpenAI says it applies stronger visual judgment to the sites, games, and renderings it builds than earlier models did.
Better judgment on ambiguous instructions
When instructions leave room for interpretation, GPT-6 Astra is designed to make better calls than earlier models, using context to fill in routine gaps and asking focused questions only when the answer could meaningfully change the outcome. In Codex, it can ask a clarifying question asynchronously while continuing other work that does not depend on the answer, proceeding with sensible assumptions if the user does not respond, but waiting on genuinely consequential decisions. OpenAI also says Astra is better at staying oriented as a task evolves, incorporating new requirements or steering messages without losing track of the original request.
Also Read >>>> ChatGPT Ads Launch in India: Who Sees Them, Who Doesn’t
Coding and long running agentic work
OpenAI describes GPT-6 Astra as its best model for software engineering to date. On Terminal-Bench 4.0, Astra scores 57.9 percent against Sol’s 37.3 percent, and on the internal Database Migration Tasks evaluation it reaches 63.9 percent versus Sol’s 42.7 percent. Alongside Astra, OpenAI is introducing a new way for Codex to preserve context as a session runs long: instead of relying only on repeated compaction and summarization, which can drop details about why a fix failed or how a component behaves, Codex can now keep searchable notes across context windows, so it can retrieve a requirement or test result from earlier in a session even if that detail was never captured in its own notes. This feature is currently opt in through a Codex configuration file and is expected to become the default for Astra in the coming weeks.
Advancing scientific discovery
OpenAI says GPT-6 Astra helped produce two new mathematical results on the gaps between prime numbers, including tightening a decades old bound on how close together certain prime pairs can occur down to 186, and improving a separate bound on unusually large prime gaps that had not moved in more than 80 years. On GPQA Diamond, a graduate level science benchmark, Astra scores 96.0 percent, ahead of Gemini 3.8 Flash’s 95.3 percent, Sol’s 94.6 percent, and Claude Fable 5.1’s 93.7 percent.
Cybersecurity: a capability jump that comes heavily gated
This is the part of the Astra story worth paying closest attention to. GPT-6 Astra is a significant jump in cyber capability, and OpenAI confirms it meets the Critical threshold in cybersecurity under its Preparedness Framework, making it the first OpenAI model to do so. Tested without production safeguards, Astra achieved a perfect 100 percent on ExploitBench, which evaluates whether a model can turn a known software vulnerability into a working exploit, compared with 78.5 percent for GPT-5.6 Sol. On a newer, contamination controlled version of that benchmark built from vulnerabilities disclosed only in the three months before launch, Astra scored far higher than Sol’s 5.5 percent, and during testing it discovered and used two previously unknown zero day vulnerabilities, which OpenAI says it is disclosing responsibly to their maintainers. On SRE-Bench, which measures whether a model can reverse engineer a software binary without access to source code, Astra solved 88.0 percent of tasks on a single attempt and 99.2 percent within four attempts, against 55.9 percent and 68.7 percent for Sol.
Because of this jump, OpenAI is shipping Astra’s cyber capabilities heavily gated. At launch, the model will help with defensive work such as secure code review and patching, but it refuses more advanced tasks such as creating proof of concept exploits. OpenAI plans to expand access to less restricted safeguards through its Daybreak program in the coming weeks, aimed at enabling more defensive workflows such as vulnerability validation, malware analysis, and detection engineering. Anyone planning to build security tooling on Astra should budget for these safeguards occasionally pausing or interrupting legitimate work while OpenAI continues to tune the system.
Alignment and safety
OpenAI calls Astra its most aligned model yet. In an internal evaluation modeled on the Hugging Face incident, which tests whether a model facing a difficult or impossible task will exceed its intended scope, GPT-5.6 Sol did so 48 percent of the time when run without production safeguards, while Astra did so in 0 percent of cases. Astra also never attempted to circumvent a Codex Auto Review denial in internal testing, even when that review step was deliberately made easy to evade and the task was otherwise impossible to complete, and OpenAI reports it is roughly three times less likely than Sol to make inaccurate claims about its own capabilities. One caveat OpenAI itself discloses: Astra’s written reasoning was found to be harder to monitor than Sol’s on tests that explicitly asked it to evade monitoring, something the company says it takes seriously and is continuing to research.
GPT-6 Astra Benchmarks: How It Stacks Up
Independent of OpenAI’s own framing, a few things stand out once the benchmark tables are read closely, and they matter for anyone deciding whether to build on Astra.
The marquee number in OpenAI’s launch materials is a 99.9 percent score on ARC-AGI-3, a benchmark specifically designed to stay ahead of AI capability. That score, however, was produced using OpenAI’s own stateful provider adapter harness. According to analysis that cross references OpenAI’s launch tables with independent runs from the ARC Prize Foundation, a standard stateless API call scores only around 17 to 63 percent on the same benchmark, depending on reasoning tier. Developers calling the model through the ordinary API, in other words, should not expect anywhere near 99 percent out of the box.
It is also worth knowing that GPT-6 Astra does not top every chart. On Humanity’s Last Exam with tools, Astra scores 57.2 percent, actually trailing Claude Fable 5.1’s 65.0 percent, Claude Fable 5’s 63.8 percent, and Claude Opus 5’s 63.6 percent. On the third party Artificial Analysis Intelligence Index, Astra’s 61.2 sits behind Claude Fable 5.1’s 65.7 and Claude Opus 5’s 63.1. And on a couple of narrower coding evaluations, FrontierCode 1.1 Extended and Main, Claude Fable 5 and Claude Opus 5 edge Astra out by a fraction of a percentage point. None of this erases Astra’s genuinely strong results elsewhere, but it does mean the honest summary is a model with a clear specialization in agentic execution and computer use, rather than a clean sweep of every metric.
Here is a snapshot of how GPT-6 Astra compares with GPT-5.6 Sol, Claude Fable 5.1, Claude Opus 5, and Gemini 3.8 Flash across some of the most cited benchmarks.
Computer use: on OSWorld 2.0, Astra scores 72.6 percent versus Sol’s 65.7 percent and Claude Opus 5’s 70.2 percent. On Agents’ Last Exam, Astra reaches 59.3 percent, ahead of Claude Opus 5’s 55.5 percent and Sol’s 53.6 percent, while OpenAI says it uses roughly 65 percent fewer output tokens than Opus 5 to get there.
Math and science: on FrontierMath Tier 4, Astra scores 97.6 percent, described in OpenAI’s own summary as effectively saturating the benchmark, ahead of Claude Fable 5.1’s 87.8 percent, Sol’s 83.0 percent, and Claude Opus 5’s 73.2 percent.
Coding: on Terminal-Bench 4.0, Astra scores 57.9 percent against Claude Fable 5.1’s 55.8 percent, Claude Opus 5’s 52.6 percent, Sol’s 37.3 percent, and Gemini 3.8 Flash’s 19.1 percent.
Cybersecurity: on ExploitBench, Astra’s 100 percent leads Sol’s 78.5 percent and Claude Opus 5’s 70 percent by a wide margin.
Long context: on an internal long context retrieval test spanning 512K to 1M tokens, Astra scores 96.3 percent against Sol’s 73.8 percent, consistent with its up to 1 million token context window.
GPT-6 Astra Pricing and Availability
Standard API pricing for GPT-6 Astra is 10 dollars per million input tokens and 50 dollars per million output tokens, with separate rates applying to cache reads and writes. OpenAI also offers a faster processing option in the API, priced at twice the Standard rate, which in practice works out to roughly 20 dollars per million input tokens and 100 dollars per million output tokens. For context, that pricing sits well above GPT-5.6 Terra’s 2 dollar and 12 dollar rates and Claude Opus 5’s 5 dollar and 25 dollar rates, positioning Astra as a premium reasoning and automation model rather than a bulk text workhorse.
In production, actual pricing and performance vary a little by provider. Based on data from OpenRouter, which routes requests across multiple hosts of the model, a discounted OpenAI Flex tier prices the model at around 5 dollars and 25 dollars per million tokens with somewhat lower throughput, while standard OpenAI and Azure endpoints price it at 10 dollars and 50 dollars, and a faster OpenAI tier reaches roughly 20 dollars and 100 dollars per million tokens. Across providers, GPT-6 Astra’s best observed throughput is around 51 tokens per second with roughly 3.75 seconds of latency at the median.
GPT-6 Astra usage is included within existing ChatGPT subscription allowances, with the option for users and businesses to purchase additional credits for extra usage. In terms of real world adoption, data from OpenRouter shows the model already being used heavily inside agent focused applications shortly after launch, with the most traffic coming from Hermes Agent, Codex, Cursor, and other coding and automation tools.
GPT-6 Astra vs Claude and Gemini
Astra arrives into a genuinely crowded frontier tier. Anthropic’s Claude Fable 5.1, released just days earlier on September 1, 2026, and Claude Opus 5 are the most frequent reference points in OpenAI’s own comparison charts, alongside Google’s Gemini 3.8 Flash. Broadly, Astra leads on computer use, most cybersecurity evaluations, and several math and science benchmarks, while Claude Fable 5.1 leads on Humanity’s Last Exam and the broader Artificial Analysis Intelligence Index, and Claude Opus 5 and Claude Fable 5 stay competitive on a handful of narrower coding evaluations. Anyone choosing between these models for a specific workload is likely better served comparing scores on the benchmark closest to that actual use case than relying on any single headline claim.
Final Take
GPT-6 Astra is a genuine step forward for OpenAI, particularly in computer use, professional document creation, and long running agentic coding, and its alignment testing results suggest real, measurable safety progress over GPT-5.6 Sol. At the same time, its most eye catching number, the 99.9 percent ARC-AGI-3 score, depends on an expensive, stateful harness that ordinary API calls will not replicate out of the box, and the model does not lead every benchmark it appears on. Its jump into Critical level cybersecurity capability is arguably the most important part of this launch to watch, since it comes with meaningful new safeguards that OpenAI is still tuning, and that can pause even legitimate defensive security work in the meantime. Anyone planning to build on Astra, especially for security adjacent workflows, should read OpenAI’s own system card before committing.
Frequently Asked Questions
What is GPT-6 Astra?
GPT-6 Astra is OpenAI’s new frontier flagship AI model, released on September 5, 2026, succeeding GPT-5.6 Sol. OpenAI describes it as its most intelligent and aligned model, built around state of the art computer use, professional knowledge work, and a major jump in cybersecurity capability.
When was GPT-6 Astra released and how can I access it?
GPT-6 Astra was released on September 5, 2026, rolling out first to a limited set of organizations and then to all ChatGPT Plus, Pro, Business, and Enterprise users. It is also available through the OpenAI API as gpt-6-astra, as well as through Microsoft Azure and AWS Bedrock.
How much does GPT-6 Astra cost to use through the API?
Standard API pricing for GPT-6 Astra is 10 dollars per million input tokens and 50 dollars per million output tokens, with a faster processing mode available at twice that price. Separate rates apply to cache reads and writes, and pricing can vary slightly depending on which provider serves the request.
What is the context window of GPT-6 Astra?
GPT-6 Astra supports a context window of up to 1 million tokens, and OpenAI reports strong retrieval accuracy even when relevant information sits between 512,000 and 1 million tokens into the context.
Is GPT-6 Astra really as good as OpenAI’s 99.9 percent ARC-AGI-3 score suggests?
That score was achieved using OpenAI’s own stateful provider harness. Independent analysis suggests that standard, stateless API calls score meaningfully lower, in the range of roughly 17 to 63 percent depending on reasoning tier, so developers should not expect the headline figure out of the box.
How does GPT-6 Astra compare to Claude Fable 5.1 and Claude Opus 5?
GPT-6 Astra generally leads on computer use, most cybersecurity benchmarks, and several math and science evaluations, while Claude Fable 5.1 scores higher on Humanity’s Last Exam and a broad third party intelligence index, and Claude Opus 5 stays competitive on a few narrower coding benchmarks. Neither model sweeps every category.
Why is GPT-6 Astra’s cybersecurity capability being restricted?
GPT-6 Astra is the first OpenAI model to cross the Critical threshold for cybersecurity capability under the company’s Preparedness Framework. As a result, it refuses advanced tasks such as creating proof of concept exploits at launch, with OpenAI planning to expand access to less restricted safeguards for defensive security work through its Daybreak program over time.
What is GPT-6 Astra Pro?
GPT-6 Astra Pro is the same underlying GPT-6 Astra model served with its reasoning mode set for higher quality responses on complex tasks. It is available to users on OpenAI’s Pro, Business, and Enterprise plans.
