GPT-6 Astra
- Provider
- OpenAI
- Status
- Current
- Context
- 1,000,000 tok
- Price
- $10 / $50 /MTok
- Knowledge
- 2026-04-30
GPT-6 Astra is OpenAI’s flagship generational release, which began rolling out on 3 September 2026 — one day after OpenAI rated it the first of its models to cross the Critical cybersecurity threshold in its own Preparedness Framework. It is priced at $10/$50 per million tokens, 2.5x GPT-5.6 Sol’s current rate and identical to Claude Fable 5.1’s.
The independent measurement puts it level with the model it replaces on general intelligence and well ahead of it on agentic coding: Artificial Analysis scores Astra 61 on its Intelligence Index — tied with Sol and Grok 4.6, below Claude Opus 5 (63), Fable 5 (62) and Fable 5.1 (66) — but 67 on its Coding Agent Index, equalling Opus 5 and Fable 5 at less than half Fable 5’s cost per task. That coding-agent result, not the headline index, is the case for the model.
The cyber capabilities that earned the Critical rating — a perfect 100% on ExploitBench, two zero-day vulnerabilities found unaided in evaluation — are not in the box you buy: they are gated to vetted organisations in OpenAI’s Daybreak Blue program. What ships generally is the same model with those capabilities withheld.
Quick specs
| Provider | OpenAI |
| Tier | Flagship (GPT-6 generation) |
| Released | 3 September 2026 — rolling out |
| Status | Daybreak enterprise first; Plus, Pro, Business, Enterprise, API and AWS to follow |
| Context window | 1,000,000 tokens |
| Knowledge cutoff | 30 April 2026 |
| Input price | $10.00 / MTok ($20.00 Fast mode) |
| Output price | $50.00 / MTok ($100.00 Fast mode) |
| AA Intelligence Index | 61 (max) — level with GPT-5.6 Sol; Fable 5.1 leads at 66 |
| AA Coding Agent Index | 67 — equal with Claude Opus 5 and Fable 5 |
| Terminal-Bench 4.0 | 57.7% (Fable 5.1: 55.8%) |
| SWE-bench | Not published — no figure of any kind |
| Best for | Agentic coding at scale, computer use, long-context retrieval |
| Limitations | 2.5x Sol’s price and ~75% more per task; intelligence index flat on Sol; cyber capability gated |
What actually changed
Agentic coding is the real generation jump. On Artificial Analysis’s independent Coding Agent Index, Astra scores 67 — the same as Claude Opus 5 and Fable 5 — at less than half Fable 5’s per-task cost. AA measures it using one-third the tokens of Sol (max) and one-fifth of Opus 5’s on coding benchmarks, a roughly 70% token-efficiency gain generation-on-generation. OpenAI’s own table agrees on direction: Terminal-Bench 4.0 rises to 57.7% (past Fable 5.1’s 55.8%), AutomationBench more than doubles Sol (41.4% vs 18.1%).
Computer use moves furthest. ScreenSpot-Pro jumps from Sol’s 76.9% to 92.7%, OSWorld 2.0 from 65.7% to 72.6%, and Agents’ Last Exam from 53.6% to 59.3% — all past the Anthropic frontier models on OpenAI’s table.
General intelligence does not move. The independent index reads 61 — the same as Sol. AA also measures regressions: roughly 80 Elo down on GDPval-AA v2 (knowledge work) versus Sol, with small declines on banking support, scientific coding and long-context reasoning tasks. On OpenAI’s own table Astra trails Fable 5.1 on Humanity’s Last Exam with tools, 57.2% to 65.0%. Astra is a specialist generation — agents, computer use, cyber — not a general-intelligence leap.
Hallucination roughly halves. On AA’s hallucination evaluation Astra’s rate falls to 51% against Sol’s 92% — the largest safety-adjacent improvement in the independent data.
The cyber capabilities are the headline OpenAI wrote for itself. Astra is the first OpenAI model rated Critical for cybersecurity under its Preparedness Framework: a perfect score on ExploitBench (turning known vulnerabilities into working exploits), two zero-days found unaided in a separate evaluation, and 88.0% on SRE-Bench against Opus 5’s 12.5%. OpenAI reports the model declines 91.5% of cyber-related jailbreak attempts against Sol’s 59%. Full cyber capability goes only to the Daybreak Blue program.
Benchmark performance
All figures OpenAI-reported and vendor-run from the 3 September launch table, except where marked independent.
Coding and agents
| Benchmark | Astra | Fable 5.1 | Opus 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 57.7% | 55.8% | 52.3% | 37.3% |
| DeepSWE v1.1 | 74.1% | 67.4% | 73.7% | — |
| Terminal-Bench-Science 0.1 | 64.6% | 52.6% | 30.0% | 22.4% |
| Agents’ Last Exam | 59.3% | 48.7% | 55.5% | 53.6% |
| AutomationBench | 41.4% | — | — | 18.1% |
The DeepSWE row deserves its caveat: Astra’s 74.1% leads Gemini 3.8 Flash (73.8%) and Opus 5 (73.7%) by fractions of a point on OpenAI’s own table, and Meta’s launch table the same day claims 75.4% for the preview-only Muse Spark 1.3 max — four vendors’ numbers within 1.7 points, none independently adjudicated.
Computer use and long context
| Benchmark | Astra | GPT-5.6 Sol | Best Anthropic |
|---|---|---|---|
| ScreenSpot-Pro | 92.7% | 76.9% | 87.3% (Fable 5) |
| OSWorld 2.0 | 72.6% | 65.7% | 70.2% (Opus 5) |
| BrowseComp | 91.5% | 90.4% | 90.8% (Opus 5) |
| MRCR v2 8-needle, 256–512K | 100.0% | 91.5% | — |
| MRCR v2 8-needle, 512K–1M | 96.3% | 73.8% | — |
The independent numbers
Artificial Analysis scores Astra (max) at 61 on the Intelligence Index — 8th of the roughly 200 models it tracks, level with Sol and Grok 4.6 — and 67 on the Coding Agent Index, equal with Opus 5 and Fable 5 at less than half Fable 5’s per-task cost. The efficiency gain is real (a third of Sol’s tokens on coding), and so is the price problem: at max effort AA measures Astra about 75% more expensive per completed task than Sol, because the 2.5x price rise outruns the token savings. The non-reasoning configuration scores 55.
What OpenAI did not publish
No SWE-bench figure of any kind — the second frontier launch in a week to skip the industry’s most-watched coding benchmark, after Fable 5.1 did the same. The highest SWE-bench Pro score on our board remains Claude Fable 5’s 80.3%, a June number. Also unpublished at launch: max output tokens, training cutoff, and any output-speed figure — AA lists throughput as not yet measured.
The Critical cyber rating, and what it means for buyers
Astra crossed the Critical threshold of OpenAI’s Preparedness Framework on cybersecurity — the first OpenAI model to do so, announced 2 September, the day before rollout began. The evidence OpenAI cites: a perfect score on ExploitBench and two zero-day vulnerabilities discovered without human intervention in a separate evaluation against recently disclosed flaws.
Practically, this splits the product in two:
- The general release ships with those capabilities withheld. OpenAI says its safeguards “sufficiently minimize the risk of severe harm” for release, and reports 91.5% refusal of cyber jailbreak attempts in testing, up from Sol’s 59%.
- The full cyber capability goes only to vetted organisations in the Daybreak Blue program — the same week Google gated Gemini 3.8 Flash Cyber behind its Fairwind Program, and a month after OpenAI gated GPT-5.6-Cyber behind Daybreak Red. Frontier cyber capability is now uniformly access-controlled; the question labs are deciding is who qualifies, not whether to build it.
How to access GPT-6 Astra
Rollout began 3 September 2026 with enterprise customers in the Daybreak programme. ChatGPT Plus, Pro, Business and Enterprise, the API and AWS follow over the coming days, per OpenAI. Fast mode — the same model at higher throughput — is priced at double the standard rate ($20/$100).
At the time of writing OpenAI has not published maximum output length or throughput figures, and API availability is still propagating; check OpenAI’s pricing page for the current state before committing a workload.
How GPT-6 Astra compares
vs GPT-5.6 Sol
Same intelligence-index score (61), very different profile. Astra is far stronger on agents and computer use (ScreenSpot-Pro 92.7% vs 76.9%, AutomationBench 41.4% vs 18.1%), dramatically better on long context, and hallucinates roughly half as often on AA’s measurement — but it costs 2.5x Sol’s promotional rate and about 75% more per completed task, and it regresses on knowledge work (~80 Elo on GDPval-AA v2). Sol stays on sale at $4/$20 to at least 21 November. If your workload is chat, drafting or analysis, Sol is the better buy; Astra earns its premium on agentic and computer-use work.
vs Claude Fable 5.1
Identical pricing ($10/$50). Fable 5.1 leads clearly on the independent intelligence index (66 vs 61) and on Humanity’s Last Exam with tools (65.0% vs 57.2% — on OpenAI’s own table). Astra leads on Terminal-Bench 4.0 (57.7% vs 55.8%), Terminal-Bench-Science (64.6% vs 52.6%), FrontierMath Tier 4 (97.6% vs 87.8%) and everything computer-use. Neither published a SWE-bench figure. On AA’s Coding Agent Index they are not equal in cost: Astra matches Fable 5’s score at less than half its per-task cost.
vs Claude Opus 5
Opus 5 at $5/$25 remains the value anchor of the frontier: a higher intelligence index (63 vs 61), the same Coding Agent Index (67), a published SWE-bench Pro figure (79.2%) and the lowest per-task cost of the frontier models. Astra beats it on computer use and long context. For most buyers Opus 5 is still the rational default; see best AI models for the full board.
vs Gemini 3.8 Flash
The odd couple of the week: Gemini 3.8 Flash matches Astra to within half a point on DeepSWE (73.8% vs 74.1%, OpenAI’s table) at $0.75/$3.75 — one-thirteenth of Astra’s price. Astra is broadly more capable (AA 61 vs 59) and far stronger on computer use; the Flash undercuts it catastrophically on cost for agentic coding workloads that fit its profile.
Known limitations
No SWE-bench figure of any kind. The pattern from Fable 5.1’s launch repeats. On coding-weighted comparisons the burden of proof sits on vendor benchmarks alone.
2.5x the price for the same measured intelligence. AA’s index reads 61 for both Astra and Sol; the per-task cost is about 75% higher at max effort.
Knowledge work regresses. ~80 Elo down on GDPval-AA v2 versus Sol, with small declines on banking support, scientific coding and long-context reasoning tasks in AA’s suite.
Still rolling out. Consumer tiers, the API and AWS were “coming days” at launch; max output and throughput are unpublished.
The headline cyber capability is not purchasable. Critical-rated capabilities live behind Daybreak Blue vetting; the general model is deliberately weaker than the one in the announcement.
Vendor-run headline benchmarks. Every figure except the Artificial Analysis measurements is OpenAI’s own.
Version history
| Version | Released | Key points |
|---|---|---|
| GPT-6 Astra | 3 Sep 2026 | First OpenAI model rated Critical for cyber; AA II 61, Coding Agent Index 67; $10/$50; no SWE-bench figure |
| GPT-5.6 (Sol, Terra, Luna) | 30 Jul 2026 | Three-tier flagship; Sol AA II 61; promotional $4/$20 to 21 Nov 2026 |
| GPT-5.5 | 2026 | Previous flagship; AA II 55 |
Frequently asked questions
Is GPT-6 Astra the best AI model?
No. On the independent Artificial Analysis Intelligence Index it scores 61 — level with GPT-5.6 Sol and Grok 4.6, and behind Claude Fable 5.1 (66), Opus 5 (63) and Fable 5 (62). We rank it #6 on our board: its Coding Agent Index score of 67 equals the Anthropic frontier at a lower per-task cost, and its computer-use results lead the field, but the general-intelligence measurement did not move from Sol.
How much does GPT-6 Astra cost?
$10 per million input tokens and $50 per million output — 2.5x GPT-5.6 Sol’s current promotional rate and identical to Claude Fable 5.1. Fast mode is double that, at $20/$100. Cache reads carry a 90% discount. Artificial Analysis measures it at roughly 75% more per completed task than Sol at max effort despite much better token efficiency.
Why was GPT-6 Astra rated a Critical cyber risk?
It is the first OpenAI model to cross the Critical cybersecurity threshold in OpenAI’s own Preparedness Framework: it scored 100% on ExploitBench, which measures turning known vulnerabilities into working exploits, and found two zero-day vulnerabilities unaided in a separate evaluation. Those capabilities are withheld from the general release and gated to vetted organisations in OpenAI’s Daybreak Blue program.
What is GPT-6 Astra’s SWE-bench score?
OpenAI did not publish one — no SWE-bench Pro or Verified figure of any kind appears in the launch material, making Astra the second frontier launch in a week (after Claude Fable 5.1) to omit it. The nearest published coding-agent evidence is vendor DeepSWE v1.1 (74.1%) and the independent AA Coding Agent Index (67).
What is the context window?
1,000,000 tokens, with a knowledge cutoff of 30 April 2026 per Artificial Analysis. OpenAI’s long-context retrieval numbers are strong: 100% on its MRCR v2 8-needle test at 256–512K and 96.3% at 512K–1M, against Sol’s 91.5% and 73.8%.
Should I upgrade from GPT-5.6 Sol?
Only if your workload is agentic. Astra’s gains are concentrated in coding agents, computer use and long context; its measured general intelligence is identical to Sol’s, it regresses on knowledge work, and it costs about 75% more per completed task. Sol remains on sale at $4/$20 until at least 21 November 2026.
Last verified 4 September 2026, the day after a rollout that is still propagating. Benchmark figures are OpenAI-reported from the 3 September 2026 launch table and vendor-run unless marked independent; Intelligence Index, Coding Agent Index, token-efficiency, cost-per-task and hallucination figures are Artificial Analysis measurements. OpenAI published no SWE-bench figure for this model. Pricing and availability subject to change.