GPT-6 Astra: Everything You Need to Know About OpenAI’s New Model
Last Updated:
2026-09-04
GPT-6 Astra: Everything You Need to Know About OpenAI’s New Model
It’s finally here—OpenAI released GPT-6 Astra on September 3, 2026, as its new flagship reasoning model for long, complex, multi-step work across coding, browser and computer use, professional analysis, science, mathematics, and healthcare.
GPT-6 Astra is already available through the OpenAI API and is gradually rolling out across ChatGPT, Codex, and selected platforms.
In this guide, we cover everything you need to know about GPT-6 Astra following its release—from technical specifications and benchmark results to API access and where the model is available now.
GPT-6 Astra is a closed multimodal reasoning model released by OpenAI on September 3, 2026. Its main difference from GPT-5.6 Sol is noticeably stronger and more efficient agentic work.
On the independent Artificial Analysis Intelligence Index, Astra has the same score as Sol and is five points behind Fable 5.1, suggesting that the gains are localized to math, tool use, and reasoning.
With that performance context in mind, it helps to clarify exactly what OpenAI released and how the company positions the model.
If there was any confusion before, we now know for sure that Astra is indeed GPT-6. Before release, “Astra” was used as the name of the next major model, which led to early speculation that it might be GPT-5.7 or a separate model family. After the official launch, however, there is no longer any ambiguity—GPT-6 is officially here.
OpenAI’s API guidance describes several new Astra features that developers will appreciate:
Asynchronous tool calls. The model can continue reasoning or work on an independent part of a task while the application executes a tool.
Mid-turn steering. The user can change requirements during execution; with a WebSocket connection, work already completed is preserved.
Changing reasoning effort without resetting the cache.configuration_update changes the reasoning effort within an ongoing conversation.
Long-term memory in Codex. An experimental notes and retrieval mechanism across earlier context windows reduces information loss from repeated compaction.
Production monitoring. Astra’s tool-use trajectories are checked asynchronously for potentially unauthorized actions.
Those features only matter once you can actually reach the model, and the rollout is still in progress.
Where to Access GPT Astra
As of September 4, 2026, GPT Astra is in a gradual rollout across ChatGPT products—to see if it is available to you, log in to your ChatGPT account or OpenAI’s developer console. In general, Astra is already being rolled out to:
Standard ChatGPT Chat: Astra appears as GPT-6 Pro. Access is announced for the $100 and $200 Pro plans, Business, and Enterprise. Plus does not receive GPT-6 Pro in standard Chat.
ChatGPT Work and Codex: Astra is gradually becoming available to Plus, Pro, Business, and Enterprise users. Plus and Business Standard receive limited usage; Pro and Business Premium use their existing shared allowance.
Enterprise: An administrator must enable the model separately. It is disabled by default at launch, and previous Early Model Access settings do not carry over automatically.
API: The model ID is gpt-6-astra; the Free API tier is not supported.
Cloud platforms:Microsoft Foundry/Azure and Amazon Bedrock have announced support. Microsoft began its rollout through a Limited Access Program.
Astra is not included in ChatGPTFree and Go plans.
GPT-6 Pro limits in Chat are separate from Work and Codex limits and may change. At the time of this report, Pro $200 includes 200 messages per week, Pro $100 has a shared limit of 50 messages per week together with GPT-5.6 Sol Pro, and Business Standard includes 15 messages per month. Business Premium includes 50 per week.
For developers and teams, availability is only half of the decision. Astra’s higher per-token price determines which workloads can justify using it.
GPT-6 Astra API Pricing and Rate Limits
According to the official GPT-6 Astra API documentation, Astra is 2.5× more expensive per token than GPT-5.6 Sol, which costs $4 per million input tokens and $20 per million output tokens. Here is the full pricing table:
Component
Price per 1 million tokens
Input
$10.00
Cached input
$1.00
Cache write
$12.50
Output
$50.00
When input exceeds 272,000 tokens, the entire request is billed at 2× the standard rate for input and cache, and 1.5× for output. Batch and Flex cost 50% of Standard. API Fast mode costs 2× and promises up to 2× the speed; in Work and Codex, the listed Fast mode credit multiplier is 2.5×. Fast mode is unavailable for Astra with EU data residency.
Published API rate limits are as follows:
Tier
Requests per minute
Tokens per minute
Batch queue limit
Tier 1
500
500,000
1,500,000
Tier 2
5,000
1,000,000
3,000,000
Tier 3
5,000
2,000,000
100,000,000
Tier 4
10,000
4,000,000
200,000,000
Tier 5
15,000
40,000,000
15,000,000,000
The price sets the bar. The benchmarks show how close Astra comes to clearing it.
How Does GPT-6 Astra Perform in Benchmarks?
Is GPT Astra better than Sol or Fable? In some cases it is, while in others it performs similarly or even slightly worse. Let’s look at GPT Astra performance across coding, computer use, science, and more.
GPT-6 Astra vs GPT-5.6 Sol benchmark comparison
Agentic Coding
In OpenAI’s official evaluations, Astra scored 57.9% on Terminal-Bench 4.0, compared with 37.3% for GPT-5.6 Sol, and 74.1% on DeepSWE v1.1, compared with 72.7%. On FrontierCode 1.1 Extended, it scored 64.5% versus 60.6%; on the Main version, 53.3% versus 47.5%.
Artificial Analysis independently gave Astra a score of 67 on its Coding Agent Index. That indicates performance similar to Claude Opus 5 and Fable 5, but below Fable 5.1, which scored 70 points. However, Astra uses roughly one-third as many tokens as GPT-5.6 Sol in the Codex harness. At maximum reasoning effort, the cost per task was approximately the same as Sol while Astra scored two points higher.
Computer and Browser Use
OpenAI’s computer-use evaluations show another area of improvement, meaning that the model is better at automatically completing tasks on behalf of the user:
Benchmark
Astra
GPT-5.6 Sol
Agents’ Last Exam
59.3%
53.6%
OSWorld 2.0
72.6%
65.7%
ScreenSpot-Pro
92.7%
76.9%
In the OSWorld simulation, Astra completed tasks in approximately 40 minutes, compared with 75 minutes for Sol—about 47% faster. In its updated Codex harness, OpenAI reports a 1.9× speedup on Mind2Web.
Professional Work
OpenAI reports that Astra is trained to work with documents, spreadsheets, and presentations, follow templates and visual styles, and extract only relevant information from a large context.
Benchmark
Astra
GPT-5.6 Sol
AutomationBench
41.4%
18.1%
BenchCAD
95.9%
83.3%
BrowseComp
91.5%
90.4%
Internal Design Tasks
50.0%
47.4%
Internal Data Science Tasks
40.9%
30.5%
Science and Mathematics
OpenAI’s math and science benchmarks are another area where Astra pulls ahead of other frontier models:
Benchmark
Astra
GPT-5.6 Sol
Best competitor
FrontierMath Tier 4 v2
97.6%
83.0%
Fable 5.1: 87.8%
GPQA Diamond
96.0%
94.6%
Gemini 3.8 Flash: 95.3%
Terminal-Bench Science 0.1
64.6%
22.4%
Fable 5.1: 52.6%
Humanity’s Last Exam with tools
57.2%
Not reported
Fable 5.1: 65.0%
These gains are substantial, but they are concentrated in particular kinds of work. The broader comparison with GPT-5.6 Sol is more nuanced.
Is GPT Astra Better than GPT-5.6 Sol?
This depends on the application, but Astra does not perform better than Sol in all tasks. For example, Artificial Analysis gives Astra a maximum score of 61 on its Intelligence Index—the same as GPT-5.6 Sol. This is slightly below Claude Fable 5.1 and Meta Muse Spark 1.3. However, it is worth noting that Astra uses approximately 10% fewer output tokens. On balance, its per-token price is 2.5× higher than that of Sol, making the average Intelligence Index task about 75% more expensive.
In some tests, Astra even scores below GPT-5.6 Sol:
a regression of approximately 80 Elo on GDPval-AA v2;
a decline of two to three points on tau3-Banking, SciCode, and AA-LCR;
lower presentation quality than Sol in Artificial Analysis’s evaluation.
Astra makes sense for long-running agentic tasks, coding, research, and computer use. But for simple chat, classification, data extraction, and high-volume requests, there’s little reason to use it over cheaper and faster models.
That distinction becomes clearer when the benchmark results are mapped to actual workloads.
When to Use GPT-6 Astra
Astra is best suited for complex, long-running tasks, especially those that are difficult to recover from if the model makes a mistake. For example, it is a strong candidate for:
Large refactors.
Complex analysis.
Agentic programming.
Application and tool control.
Creation of complete document sets.
Work with very large contexts.
Defensive-security tasks where errors carry meaningful risk.
Using Astra is probably overkill for simple question answering, classification, extraction, routing, short code generation, and routine support work—all of which will usually be better suited to GPT-5.6 Terra, GPT-5.6 Luna, or another lower-cost model.
Teams that decide Astra fits their workload should also account for several differences in API configuration and prompt design.
How to Configure and Prompt GPT-6 Astra
If you’re migrating to Astra from GPT-5.6, OpenAI recommends making a few changes to both API configuration and prompt design.
API Configuration
Astra does not support reasoning.effort: none; the lowest available setting is low. If an existing integration uses none or minimal, OpenAI recommends starting with low and increasing the effort when the task benefits from additional reasoning.
Some sampling controls should also be removed. Astra does not support temperature, top_p, or top_logprobs, and logprobs is unsupported in Chat Completions.
For tool-using applications, use the Responses API rather than building new integrations around Chat Completions.
Prompt Design
Astra is generally better at following long and detailed instructions, but that also makes contradictory instructions more consequential. Before migrating an agent, review AGENTS.md, skills, system instructions, repository guidance, and other injected context for duplicate or conflicting rules. Astra may follow conflicts more literally rather than quietly ignoring one of them.
It is also worth being more explicit about the desired output. Astra tends toward detailed, structured, and heavily formatted responses unless told otherwise. If you want a short answer, plain prose, a fixed schema, or no headings, say so directly.
For example:
“Answer in no more than five sentences.”
“Use plain prose with no headings or bullet points.”
“Return only the requested JSON object.”
“Make the smallest necessary code change.”
Astra is also more willing to ask a clarifying question when missing information could materially affect the result. For interactive work this is often useful. For automated workflows, however, it may be preferable to specify what the model should do when information is ambiguous—for example, make the safest reasonable assumption, preserve existing behavior, or stop rather than guess.
Coding Agents
On small coding tasks, Astra can sometimes verify more broadly than necessary: running large test suites, inspecting unrelated files, or performing additional checks after the requested change is already complete.
If latency or tool usage matters, constrain the scope explicitly. For example:
“Run only tests directly related to the changed module.”
“Do not run the full test suite unless the targeted tests fail.”
“Inspect only files required to make this change.”
“Do not refactor unrelated code.”
To summarize, Astra benefits from clearer boundaries.
Those boundaries matter not only for cost and reliability, but also because Astra’s stronger tool use changes its security profile.
GPT-6 Astra and Cybersecurity
For most users, the Critical rating does not mean Astra will try to attack systems. It explains why some cybersecurity requests are restricted, why advanced access is gated, and why tool-using sessions may receive additional monitoring. According to the GPT-6 Astra System Card, Astra is OpenAI’s first model classified as Critical in cybersecurity capability. In evaluations without production safeguards, it:
scored 100% on ExploitBench even at the lowest verified reasoning effort;
reached 42.4% on ExploitGym, compared with 30.3% for Sol;
solved 88.0% of SRE-Bench tasks on the first attempt and 99.2% within four attempts;
found and exploited two previously unknown zero-day vulnerabilities in an internal set of recent vulnerabilities created after its knowledge cutoff;
built a browser exploit chain with sandbox escape in a laboratory evaluation with web access and up to 64 subagents, then adapted it to a stable release in 41 total hours;
created a local privilege-escalation exploit for a hardened operating-system kernel in 12 hours;
solved 86 of 226 FrontierCyber tasks in an external Irregular evaluation, compared with 34 for Sol, but solved none of the seven Elite tasks and did not successfully attack fully hardened targets.
Without guardrails, this capability profile would make Astra unusually powerful for offensive security work.
In fairness, the ExploitBench result may be inflated because older vulnerabilities appeared in the training data, and in specific cases the model may have directly recalled a CVE. Nevertheless, the broader evaluations explain why Astra’s strongest cybersecurity capabilities are not exposed without safeguards. In its standard configuration, Astra refuses requests for advanced proof-of-concept exploits, while broader defensive access is available to verified organizations through Daybreak Blue.
How Reliably Does GPT-6 Astra Follow Instructions?
For users, alignment affects whether Astra follows instructions, resists malicious directions hidden in web pages or files, and stays within defined boundaries while using tools. On these measures, OpenAI’s safety evaluations indicate that Astra performs substantially better than GPT-5.6 Sol, especially when constraints are stated clearly.
In tests designed to see whether the model would work around Codex Auto-Review, Astra never attempted to bypass the restriction, while Sol did so in 5% of runs at maximum effort. Across more than 54,000 simulated internal Codex tasks, Astra also produced roughly half as many serious misalignment flags. Its resistance to indirect prompt injection rose from 96.23% to 99.79%, while performance on an instruction-hierarchy test reached 99.99%.
The practical takeaway is that Astra is less likely to ignore, reinterpret, or route around a rule it has been given. But the evaluations also show that clear boundaries matter.
However, there has been a regression in monitorability. Astra’s internal reasoning tends to be shorter and more compressed, which means a reviewer gets less useful information from inspecting its chain-of-thought. When Astra was explicitly instructed to conceal sandbagging, a chain-of-thought monitor detected the behavior in fewer than 11% of cases.
In one test, Astra complied with instructions to control what appeared in its chain-of-thought in 93% of examples, compared with 48% for Sol. In the UK AI Security Institute’s API evaluation, reasoning summaries were missing from as many as 80% of long cyber trajectories.
This does not mean Astra is routinely hiding secret goals or deceptive plans. However, even though it can follow rules more reliably, it has become harder to evaluate by reading its reasoning traces. For long autonomous runs, permissions, action logs, and review gates therefore matter more than reasoning summaries alone.
After the practical capabilities and deployment risks, the remaining context is how OpenAI trained the model and which technical details it still keeps private.
Astra was trained from scratch. OpenAI says its training corpus contained more than 10 trillion text tokens, more than 1 billion images, and over 1 million hours each of audio and video. The dataset is multilingual, with particularly strong English coverage alongside substantial representation of EU and other languages.
The training material came from a broad mix of sources: publicly available web data, including Common Crawl; partner and private datasets; synthetic data; code; scientific and technical documents; reference works; and material produced by employees, contractors, and other professionals. OpenAI also says that, where user settings and opt-out controls allow it, interactions with ChatGPT and Codex may contribute to training data.
To create that training corpus, GPTBot gathered accessible web content from roughly March 2013 through May 2026, subject to sites’ GPTBot robots.txt settings, while some other datasets were collected as late as August 6, 2026. Those collection dates should not be read as the model’s knowledge cutoff: OpenAI separately gives Astra a knowledge cutoff of April 30, 2026.
Synthetic training data was generated in part using GPT-5.4, GPT-5.5, GPT-5.6, and specialized internal and third-party models.
However, OpenAI does not provide Astra’s architecture, parameter count, expert count or active-parameter count, vocabulary size, exact training compute, training duration, or a breakdown showing how much of the corpus came from each data source.
Axios, citing OpenAI, has separately reported that Astra was the company’s largest training run to date and used more than 100,000 GPUs at the Stargate site in Texas.
Bottom Line
GPT-6 Astra is OpenAI’s new best AI model and a meaningful step forward for complex agentic work. However, it was tuned specifically for agentic, scientific, and long-horizon tasks, so how much benefit you will see from switching to it depends on how you plan to use it. Here are the main takeaways:
Astra is rolling out across ChatGPT products, the OpenAI API, Codex, and selected cloud platforms.
OpenAI’s benchmarks show that Astra outperforms GPT-5.6 Sol most notably in agentic coding, computer use, advanced mathematics, and multi-step professional workflows.
Astra ties Sol on the Artificial Analysis Intelligence Index and remains behind Claude Fable 5.1.
Astra costs 2.5 times more per token than GPT-5.6 Sol, making it a better fit for high-value work than routine, high-volume requests.
Astra is OpenAI’s first broadly released model with a Critical cybersecurity capability rating, so access controls and monitoring are central to its deployment.