GPT-6 Astra: Release, Benchmarks, Prices, and OpenAI's AGI Claim
GPT-6 Astra is official. OpenAI unveiled the new flagship model on September 3, 2026, announcing a significant leap in computer control, coding, mathematics, cybersecurity, and long contexts. At the same time, OpenAI President Greg Brockman's statement that Astra could mark the beginning of the AGI era is generating at least as much attention as the actual benchmark figures.
As of September 5, 2026, access is still being rolled out gradually. This article therefore separates confirmed product data from OpenAI's self-assessment: what GPT-6 Astra actually costs, for whom ChatGPT Astra will be enabled, which benchmarks are particularly strong, where competing models still lead, and why a 99.9 percent score on ARC-AGI-3 is not yet scientific proof of AGI.
Quick & brief
- Release: OpenAI introduced GPT-6 Astra on September 3, 2026. The rollout began with a limited number of organizations and is expected to reach Plus, Pro, Business, and Enterprise afterward.
- API: The model is called gpt-6-astra and is also offered via Microsoft Azure and AWS Bedrock.
- Price: In the standard API, 1 million input tokens cost $10 and 1 million output tokens cost $50. Cached input costs $1 per million tokens.
- Benchmarks: Particularly striking are 99.9% on ARC-AGI-3, 97.6% on FrontierMath Tier 4, 100% on ExploitBench, and 72.6% on OSWorld 2.0.
- Not number one everywhere: On Humanity's Last Exam with Tools, Astra achieves 57.2%, while Claude Fable 5.1 achieves 65.0% in OpenAI's own comparison table.
- AGI: Greg Brockman said he personally believes OpenAI has achieved AGI. This is an interpretation by an OpenAI executive, not an independently established scientific fact.
What is GPT-6 Astra?
GPT-6 Astra is OpenAI's new frontier model for particularly demanding end-to-end tasks. OpenAI positions it not just as a better chatbot, but as a model that can directly work with software, operate browsers and desktop interfaces, write code, perform complex research, and generate professional documents. In the official announcement, OpenAI describes Astra as its most "intelligent and aligned" model to date. This phrasing is a manufacturer's assessment and should not be confused with an independent ranking.
Technically, the large context window is particularly noteworthy. The API documentation mentions 1,050,000 tokens of context and up to 128,000 output tokens. The documented knowledge cutoff date is April 30, 2026. For reasoning, developers can choose between the levels low, medium, high, xhigh, and max. Those who want to follow the development of the previous generation can find an assessment at Zerlo for GPT-5.6 Sol and its ChatGPT rollout.
GPT-6 Astra Release: Who gets ChatGPT Astra?
The launch is not a global switch that was flipped for all accounts simultaneously on September 3. OpenAI started with a limited group of organizations and announced that GPT-6 Astra would be rolled out to ChatGPT Plus, Pro, Business, and Enterprise in the following days. Pro, Business, and Enterprise customers will also get access to GPT-6 Astra Pro. For Enterprise, Astra is deactivated by default at launch and must be enabled by administrators for the workspace.
Important: OpenAI does not mention a general release for Free or Go accounts in the Astra announcement. Therefore, those who do not see the model on September 5, 2026, may simply still be in the staggered rollout despite having a valid plan. Astra usage in ChatGPT is included in existing plan quotas, according to OpenAI; additional usage will be possible via credits. So, there is no separately announced monthly 'Astra plan'.

Source: simpleicons.org
GPT-6 Astra is not only offered directly through OpenAI. OpenAI explicitly names Microsoft Azure as an additional deployment path for developers and companies.

Source: simpleicons.org
AWS Bedrock is also part of the announced Astra rollout. For existing cloud architectures, this may be more important than a direct switch to the OpenAI API.
GPT-6 Astra Benchmarks: How big is the leap?
The strongest figures are in areas where agentic models have previously hit limits particularly quickly: abstract problem-solving, computer control, software development, and offensive security tasks. OpenAI's publication includes both external benchmarks and internal tests. Independent reproduction is naturally not yet possible, especially with internal evaluations.
| Benchmark | GPT-6 Astra | GPT-5.6 Sol | Assessment |
|---|---|---|---|
| ARC-AGI-3 | 99,9 % | 7,8 % | Extremely large leap in new interactive abstraction tasks |
| FrontierMath Tier 4 (v2) | 97,6 % | 83,0 % | Very high performance in demanding mathematics |
| Terminal-Bench 4.0 | 57,9 % | 37,3 % | Significant progress in agentic coding in the terminal |
| OSWorld 2.0 | 72,6 % | 65,7 % | Better computer control; OpenAI also reports significantly shorter task times |
| AutomationBench | 41,4 % | 18,1 % | More than double the score in professional automation tasks |
| ExploitBench | 100,0 % | 78,5 % | Shows very strong exploit development capability in controlled tests |
These figures are impressive, but they need context. OpenAI itself points out that the GPT evaluations were conducted in a research environment or via the API. Production ChatGPT may yield different results due to different system prompts, tools, and security mechanisms. Furthermore, the comparison table shows the best scores across various reasoning efforts. Astra doesn't win every test either. A good counter-example is Humanity's Last Exam with Tools: Astra achieves 57.2%, while Claude Fable 5.1 is listed at 65.0% in the same OpenAI table. The Artificial Analysis Intelligence Index for Astra is 61.2, also below several Claude models listed there. This contradicts the simplified statement that Astra is automatically the best model in every dimension.
Astra also doesn't win every test. A good counter-example is Humanity's Last Exam with Tools: Astra achieves 57.2%, while Claude Fable 5.1 achieves 65.0% in the same OpenAI table. The Artificial Analysis Intelligence Index for Astra is 61.2, also below several Claude models listed there. This contradicts the simplified statement that Astra is automatically the best model in every dimension.
Why ARC-AGI-3 gets so much attention
The leap from 7.8% for GPT-5.6 Sol to 99.9% for GPT-6 Astra is the most spectacular single figure of the launch. ARC-AGI-3 is designed to test whether a system discovers rules in new, interactive environments and acts efficiently, rather than just recalling known patterns. The ARC Prize Foundation stated in OpenAI's launch contribution that Astra surpassed the human action efficiency baseline on 96% of levels. This is a strong signal for new problem-solving capabilities – but not yet a general measure of all human abilities.
GPT-6 Astra Price: What does the API cost?
For developers, the pricing structure is clearer than the ChatGPT plans: OpenAI bills Astra on a token basis in the standard API. The values apply per 1 million text tokens.
| Billing | Price per 1 million tokens |
|---|---|
| Input | $10.00 |
| Cached Input | $1.00 |
| Cache Writes | $12.50 |
| Output | $50.00 |
With very large prompts, it gets more expensive: once a request contains more than 272,000 input tokens, OpenAI charges double the input and cache rates and 1.5 times the output rate for the entire request. According to the model documentation, Batch and Flex cost 50% of the standard rates. The Fast mode, on the other hand, costs double the applicable rate and is said to deliver up to double the speed. Thus, while the maximum context window is technically attractive, it is not automatically economical. Those who regularly feed hundreds of thousands of tokens into agent runs should take prompt caching, context selection, and task division seriously. Zerlo's older overview of ChatGPT 5.4 and its API pricing structure shows how much costs have shifted between model generations.
Has OpenAI achieved AGI with GPT-6 Astra?
The statement is confirmed – the proof is not. Axios reported from the press briefing at the launch that OpenAI President Greg Brockman personally believes OpenAI has achieved AGI with Astra, or that this model could be seen in retrospect as the arrival of Artificial General Intelligence. He also left the final classification to the users. OpenAI's official product contribution presents Astra very aggressively as a new generation of intelligence but provides no independent certification according to which an generally accepted AGI threshold would have been crossed.
The core problem is the definition. There is no universally accepted test for AGI that, upon reaching it, objectively designates a model as 'AGI'. A model can nearly saturate a single abstract benchmark and yet lag behind competing models in other areas, make errors, require safety limits, or fail in real-world long-term tasks. This is precisely what OpenAI's own benchmark table already shows: Astra is exceptionally strong on ARC-AGI-3, but not a leader on Humanity's Last Exam with Tools.
The most factually sensible formulation is therefore: GPT-6 Astra is an exceptionally strong candidate in the current AGI debate, but OpenAI's AGI claim remains a corporate and management assessment. Whether research and the public also see this threshold as reached depends on the AGI definition used and on independent replications in real-world use.
Computer Control and Agents: Where Astra Becomes Practically Relevant
The most important change compared to a classic chat model is not just more knowledge or a higher coding score, but the ability to execute tasks directly. OpenAI cites forms, CRM updates, calendar organization, research, software installation, troubleshooting, data analysis, and frontend QA as examples. On OSWorld 2.0, Astra achieves 72.6% compared to 57.2% for GPT-5.6 Sol. In OpenAI's latency simulation, Astra took approximately 40 minutes per task instead of 75 minutes – about 47% less time. For companies, this combination of reasoning, computer use, and long context can be more important than another percentage point on a knowledge benchmark. An agent that correctly understands a task, can operate multiple applications, and can maintain context over long processes can potentially transform entire workflows. At the same time, however, the importance of permissions, confirmation dialogs, logging, and clear boundaries also increases.
The downside: Astra reaches OpenAI's 'Critical' threshold for cybersecurity.
OpenAI classifies GPT-6 Astra as its first proprietary model to reach the critical threshold in the Preparedness Framework for cyber capabilities. Without production protection, Astra was able to find unknown vulnerabilities and develop exploits against hardened systems, according to OpenAI. On ExploitBench, the model achieves 100%, and on ExploitGym, 42.4%. Those who want to understand the workings of this benchmark in more detail can find a detailed explanation at Zerlo about ExploitGym and AI-powered exploit development.
OpenAI is responding to this with stronger protection mechanisms and restrictions for advanced offensive cyber tasks. At the same time, there is an unpleasant second observation: In specially adversarially constructed tests, Astra's written reasoning was more difficult to monitor than that of GPT-5.6 Sol. OpenAI attributes this, among other things, to Astra having more control over its written reasoning and being able to solve simpler problems with fewer visible intermediate steps. The company itself describes the decline in monitorability as a serious research topic.
This appears contradictory to OpenAI's statement that Astra is more "aligned" only at first glance. A model can stay within the given limits more often in normal behavior and still possess technically more dangerous capabilities. This is precisely why model behavior, base capabilities, and external protection systems must be considered separately.
What does GPT-6 Astra mean for users and developers?
For regular ChatGPT users, the most important short-term question is the rollout. Those using Plus, Pro, Business, or Enterprise should not assume that a lack of Astra access on September 5th is already an error. OpenAI explicitly states a gradual rollout over several days.
For developers, the decision is more economic. Astra is primarily worthwhile where a single successful run has high value: complex coding agents, large repositories, computer use automation, demanding research, or long professional workflows. For simple classification, short text generation, or cost-sensitive mass tasks, a smaller model may still be more sensible. The high output price of $50 per million tokens and the surcharges for very long contexts make a clean model choice more important, not less important.
For companies, a third factor comes into play: governance. A more powerful agent should not automatically be granted more rights. Especially with file systems, browsers, email, internal applications, and cloud resources, minimal permissions, confirmations for irreversible actions, and traceable audit trails should remain standard.
FAQ
When was GPT-6 Astra released?
OpenAI introduced GPT-6 Astra on September 3, 2026. The rollout began with a limited number of organizations and will subsequently be gradually expanded to ChatGPT Plus, Pro, Business, and Enterprise, as well as the API, Microsoft Azure, and AWS Bedrock.
Is GPT-6 Astra already available in ChatGPT Plus?
OpenAI has explicitly announced Plus for the rollout. Since the release is happening over several days, an eligible Plus account may still not show Astra access on September 5, 2026. This alone is not yet an indication of a technical error.
How much does GPT-6 Astra cost in the API?
The standard price is $10 per 1 million input tokens and $50 per 1 million output tokens. Cached input costs $1 and cache writes $12.50 per million tokens. Requests with more than 272,000 input tokens are subject to higher multipliers.
What is the context window size of GPT-6 Astra?
The OpenAI API documentation mentions a context window of 1,050,000 tokens and a maximum of 128,000 output tokens. However, very long requests are not only technically more demanding but also more expensive for more than 272,000 input tokens.
Has GPT-6 Astra truly achieved AGI?
This is not objectively proven. Greg Brockman said he personally believes OpenAI has achieved AGI or that Astra can be considered the beginning of the AGI era. However, there is no universally accepted scientific benchmark that defines a universal AGI threshold. The claim should therefore be read as an assessment by OpenAI and not as a concluded scientific fact.
Is GPT-6 Astra safer than GPT-5.6 Sol?
OpenAI reports better results in several alignment and boundary tests, while Astra possesses significantly stronger cyber capabilities and is more difficult to monitor via its written reasoning in adversarial tests. "Safer" therefore depends on whether one considers behavior, underlying capabilities, or monitorability.
Conclusion
GPT-6 Astra is a real and technically exceptional release, not just a rumor. The strongest confirmed points are the 1.05 million token context window, significant advances in computer use and coding, and the extremely high scores on ARC-AGI-3, FrontierMath, and cybersecurity benchmarks. At the same time, a differentiated view is still necessary: Astra is not leading in every comparison, the rollout is still ongoing, and the new cyber capability creates additional security risks.
OpenAI's AGI claim is therefore the part of the launch that requires the most interpretation. Brockman's statement is remarkable, but an almost perfect ARC-AGI-3 score alone does not turn a disputed definition into a scientifically concluded fact. For users, Astra is primarily a very powerful new ChatGPT and agent model; whether it will actually be considered the beginning of the AGI era in retrospect will only be shown through independent tests and real long-term use.