THE AI RANKINGS

OpenAI

GPT-6 Astra

Provider
OpenAI
Status
Current
Context
1,000,000 tok
Price
$10 / $50 /MTok
Knowledge
2026-04-30

GPT-6 Astra is OpenAI’s flagship generational release, which began rolling out on 3 September 2026 — one day after OpenAI rated it the first of its models to cross the Critical cybersecurity threshold in its own Preparedness Framework. It is priced at $10/$50 per million tokens, 2.5x GPT-5.6 Sol’s current rate and identical to Claude Fable 5.1’s.

The independent measurement puts it level with the model it replaces on general intelligence and well ahead of it on agentic coding: Artificial Analysis scores Astra 61 on its Intelligence Index — tied with Sol and Grok 4.6, below Claude Opus 5 (63), Fable 5 (62) and Fable 5.1 (66) — but 67 on its Coding Agent Index, equalling Opus 5 and Fable 5 at less than half Fable 5’s cost per task. That coding-agent result, not the headline index, is the case for the model.

The cyber capabilities that earned the Critical rating — a perfect 100% on ExploitBench, two zero-day vulnerabilities found unaided in evaluation — are not in the box you buy: they are gated to vetted organisations in OpenAI’s Daybreak Blue program. What ships generally is the same model with those capabilities withheld.

Quick specs

ProviderOpenAI
TierFlagship (GPT-6 generation)
Released3 September 2026 — rolling out
StatusDaybreak enterprise first; Plus, Pro, Business, Enterprise, API and AWS to follow
Context window1,000,000 tokens
Knowledge cutoff30 April 2026
Input price$10.00 / MTok ($20.00 Fast mode)
Output price$50.00 / MTok ($100.00 Fast mode)
AA Intelligence Index61 (max) — level with GPT-5.6 Sol; Fable 5.1 leads at 66
AA Coding Agent Index67 — equal with Claude Opus 5 and Fable 5
Terminal-Bench 4.057.7% (Fable 5.1: 55.8%)
SWE-benchNot published — no figure of any kind
Best forAgentic coding at scale, computer use, long-context retrieval
Limitations2.5x Sol’s price and ~75% more per task; intelligence index flat on Sol; cyber capability gated

CHECK STATUS →

What actually changed

Agentic coding is the real generation jump. On Artificial Analysis’s independent Coding Agent Index, Astra scores 67 — the same as Claude Opus 5 and Fable 5 — at less than half Fable 5’s per-task cost. AA measures it using one-third the tokens of Sol (max) and one-fifth of Opus 5’s on coding benchmarks, a roughly 70% token-efficiency gain generation-on-generation. OpenAI’s own table agrees on direction: Terminal-Bench 4.0 rises to 57.7% (past Fable 5.1’s 55.8%), AutomationBench more than doubles Sol (41.4% vs 18.1%).

Computer use moves furthest. ScreenSpot-Pro jumps from Sol’s 76.9% to 92.7%, OSWorld 2.0 from 65.7% to 72.6%, and Agents’ Last Exam from 53.6% to 59.3% — all past the Anthropic frontier models on OpenAI’s table.

General intelligence does not move. The independent index reads 61 — the same as Sol. AA also measures regressions: roughly 80 Elo down on GDPval-AA v2 (knowledge work) versus Sol, with small declines on banking support, scientific coding and long-context reasoning tasks. On OpenAI’s own table Astra trails Fable 5.1 on Humanity’s Last Exam with tools, 57.2% to 65.0%. Astra is a specialist generation — agents, computer use, cyber — not a general-intelligence leap.

Hallucination roughly halves. On AA’s hallucination evaluation Astra’s rate falls to 51% against Sol’s 92% — the largest safety-adjacent improvement in the independent data.

The cyber capabilities are the headline OpenAI wrote for itself. Astra is the first OpenAI model rated Critical for cybersecurity under its Preparedness Framework: a perfect score on ExploitBench (turning known vulnerabilities into working exploits), two zero-days found unaided in a separate evaluation, and 88.0% on SRE-Bench against Opus 5’s 12.5%. OpenAI reports the model declines 91.5% of cyber-related jailbreak attempts against Sol’s 59%. Full cyber capability goes only to the Daybreak Blue program.

Benchmark performance

All figures OpenAI-reported and vendor-run from the 3 September launch table, except where marked independent.

Coding and agents

BenchmarkAstraFable 5.1Opus 5GPT-5.6 Sol
Terminal-Bench 4.057.7%55.8%52.3%37.3%
DeepSWE v1.174.1%67.4%73.7%
Terminal-Bench-Science 0.164.6%52.6%30.0%22.4%
Agents’ Last Exam59.3%48.7%55.5%53.6%
AutomationBench41.4%18.1%

The DeepSWE row deserves its caveat: Astra’s 74.1% leads Gemini 3.8 Flash (73.8%) and Opus 5 (73.7%) by fractions of a point on OpenAI’s own table, and Meta’s launch table the same day claims 75.4% for the preview-only Muse Spark 1.3 max — four vendors’ numbers within 1.7 points, none independently adjudicated.

Computer use and long context

BenchmarkAstraGPT-5.6 SolBest Anthropic
ScreenSpot-Pro92.7%76.9%87.3% (Fable 5)
OSWorld 2.072.6%65.7%70.2% (Opus 5)
BrowseComp91.5%90.4%90.8% (Opus 5)
MRCR v2 8-needle, 256–512K100.0%91.5%
MRCR v2 8-needle, 512K–1M96.3%73.8%

The independent numbers

Artificial Analysis scores Astra (max) at 61 on the Intelligence Index — 8th of the roughly 200 models it tracks, level with Sol and Grok 4.6 — and 67 on the Coding Agent Index, equal with Opus 5 and Fable 5 at less than half Fable 5’s per-task cost. The efficiency gain is real (a third of Sol’s tokens on coding), and so is the price problem: at max effort AA measures Astra about 75% more expensive per completed task than Sol, because the 2.5x price rise outruns the token savings. The non-reasoning configuration scores 55.

What OpenAI did not publish

No SWE-bench figure of any kind — the second frontier launch in a week to skip the industry’s most-watched coding benchmark, after Fable 5.1 did the same. The highest SWE-bench Pro score on our board remains Claude Fable 5’s 80.3%, a June number. Also unpublished at launch: max output tokens, training cutoff, and any output-speed figure — AA lists throughput as not yet measured.

The Critical cyber rating, and what it means for buyers

Astra crossed the Critical threshold of OpenAI’s Preparedness Framework on cybersecurity — the first OpenAI model to do so, announced 2 September, the day before rollout began. The evidence OpenAI cites: a perfect score on ExploitBench and two zero-day vulnerabilities discovered without human intervention in a separate evaluation against recently disclosed flaws.

Practically, this splits the product in two:

How to access GPT-6 Astra

Rollout began 3 September 2026 with enterprise customers in the Daybreak programme. ChatGPT Plus, Pro, Business and Enterprise, the API and AWS follow over the coming days, per OpenAI. Fast mode — the same model at higher throughput — is priced at double the standard rate ($20/$100).

At the time of writing OpenAI has not published maximum output length or throughput figures, and API availability is still propagating; check OpenAI’s pricing page for the current state before committing a workload.

How GPT-6 Astra compares

vs GPT-5.6 Sol

Same intelligence-index score (61), very different profile. Astra is far stronger on agents and computer use (ScreenSpot-Pro 92.7% vs 76.9%, AutomationBench 41.4% vs 18.1%), dramatically better on long context, and hallucinates roughly half as often on AA’s measurement — but it costs 2.5x Sol’s promotional rate and about 75% more per completed task, and it regresses on knowledge work (~80 Elo on GDPval-AA v2). Sol stays on sale at $4/$20 to at least 21 November. If your workload is chat, drafting or analysis, Sol is the better buy; Astra earns its premium on agentic and computer-use work.

vs Claude Fable 5.1

Identical pricing ($10/$50). Fable 5.1 leads clearly on the independent intelligence index (66 vs 61) and on Humanity’s Last Exam with tools (65.0% vs 57.2% — on OpenAI’s own table). Astra leads on Terminal-Bench 4.0 (57.7% vs 55.8%), Terminal-Bench-Science (64.6% vs 52.6%), FrontierMath Tier 4 (97.6% vs 87.8%) and everything computer-use. Neither published a SWE-bench figure. On AA’s Coding Agent Index they are not equal in cost: Astra matches Fable 5’s score at less than half its per-task cost.

vs Claude Opus 5

Opus 5 at $5/$25 remains the value anchor of the frontier: a higher intelligence index (63 vs 61), the same Coding Agent Index (67), a published SWE-bench Pro figure (79.2%) and the lowest per-task cost of the frontier models. Astra beats it on computer use and long context. For most buyers Opus 5 is still the rational default; see best AI models for the full board.

vs Gemini 3.8 Flash

The odd couple of the week: Gemini 3.8 Flash matches Astra to within half a point on DeepSWE (73.8% vs 74.1%, OpenAI’s table) at $0.75/$3.75 — one-thirteenth of Astra’s price. Astra is broadly more capable (AA 61 vs 59) and far stronger on computer use; the Flash undercuts it catastrophically on cost for agentic coding workloads that fit its profile.

Known limitations

No SWE-bench figure of any kind. The pattern from Fable 5.1’s launch repeats. On coding-weighted comparisons the burden of proof sits on vendor benchmarks alone.

2.5x the price for the same measured intelligence. AA’s index reads 61 for both Astra and Sol; the per-task cost is about 75% higher at max effort.

Knowledge work regresses. ~80 Elo down on GDPval-AA v2 versus Sol, with small declines on banking support, scientific coding and long-context reasoning tasks in AA’s suite.

Still rolling out. Consumer tiers, the API and AWS were “coming days” at launch; max output and throughput are unpublished.

The headline cyber capability is not purchasable. Critical-rated capabilities live behind Daybreak Blue vetting; the general model is deliberately weaker than the one in the announcement.

Vendor-run headline benchmarks. Every figure except the Artificial Analysis measurements is OpenAI’s own.

Version history

VersionReleasedKey points
GPT-6 Astra3 Sep 2026First OpenAI model rated Critical for cyber; AA II 61, Coding Agent Index 67; $10/$50; no SWE-bench figure
GPT-5.6 (Sol, Terra, Luna)30 Jul 2026Three-tier flagship; Sol AA II 61; promotional $4/$20 to 21 Nov 2026
GPT-5.52026Previous flagship; AA II 55

Frequently asked questions

Is GPT-6 Astra the best AI model?

No. On the independent Artificial Analysis Intelligence Index it scores 61 — level with GPT-5.6 Sol and Grok 4.6, and behind Claude Fable 5.1 (66), Opus 5 (63) and Fable 5 (62). We rank it #6 on our board: its Coding Agent Index score of 67 equals the Anthropic frontier at a lower per-task cost, and its computer-use results lead the field, but the general-intelligence measurement did not move from Sol.

How much does GPT-6 Astra cost?

$10 per million input tokens and $50 per million output — 2.5x GPT-5.6 Sol’s current promotional rate and identical to Claude Fable 5.1. Fast mode is double that, at $20/$100. Cache reads carry a 90% discount. Artificial Analysis measures it at roughly 75% more per completed task than Sol at max effort despite much better token efficiency.

Why was GPT-6 Astra rated a Critical cyber risk?

It is the first OpenAI model to cross the Critical cybersecurity threshold in OpenAI’s own Preparedness Framework: it scored 100% on ExploitBench, which measures turning known vulnerabilities into working exploits, and found two zero-day vulnerabilities unaided in a separate evaluation. Those capabilities are withheld from the general release and gated to vetted organisations in OpenAI’s Daybreak Blue program.

What is GPT-6 Astra’s SWE-bench score?

OpenAI did not publish one — no SWE-bench Pro or Verified figure of any kind appears in the launch material, making Astra the second frontier launch in a week (after Claude Fable 5.1) to omit it. The nearest published coding-agent evidence is vendor DeepSWE v1.1 (74.1%) and the independent AA Coding Agent Index (67).

What is the context window?

1,000,000 tokens, with a knowledge cutoff of 30 April 2026 per Artificial Analysis. OpenAI’s long-context retrieval numbers are strong: 100% on its MRCR v2 8-needle test at 256–512K and 96.3% at 512K–1M, against Sol’s 91.5% and 73.8%.

Should I upgrade from GPT-5.6 Sol?

Only if your workload is agentic. Astra’s gains are concentrated in coding agents, computer use and long context; its measured general intelligence is identical to Sol’s, it regresses on knowledge work, and it costs about 75% more per completed task. Sol remains on sale at $4/$20 until at least 21 November 2026.


Last verified 4 September 2026, the day after a rollout that is still propagating. Benchmark figures are OpenAI-reported from the 3 September 2026 launch table and vendor-run unless marked independent; Intelligence Index, Coding Agent Index, token-efficiency, cost-per-task and hallucination figures are Artificial Analysis measurements. OpenAI published no SWE-bench figure for this model. Pricing and availability subject to change.