GPT-5.6 Sol
- Provider
- OpenAI
- Status
- Current
- Context
- 1,000,000 tok
- SWE-bench
- 64.6%
- Price
- $5 / $30 /MTok
- Knowledge
- 2026-02
GPT-5.6 is OpenAI’s current flagship model family, generally available since 9 July 2026 across ChatGPT, Codex, ChatGPT Work and the API, after a two-week government-coordinated limited preview. It ships in three tiers: Sol, the flagship; Terra, a balanced everyday model; and Luna, a fast, low-cost model (OpenAI). All three share a 1M-token context window, a 128K max output, and a February 2026 knowledge cutoff.
GPT-5.6 Sol is OpenAI’s strongest model to date. With independent numbers now in, it ranks 3rd on the Artificial Analysis Intelligence Index (59, behind Claude Opus 5 and Claude Fable 5), leads the Artificial Analysis Coding Agent Index (80), tops agentic command-line coding on Terminal-Bench 2.1, and reaches 92.5% on ARC-AGI-2 — all while being unusually token-efficient (~15k tokens per Intelligence Index task at ~$1.04, roughly a third of Fable 5’s cost). It introduces two new controls: a max reasoning effort for deeper single-model reasoning, and an ultra mode that coordinates subagents in parallel, and debuts a naming convention where the number (5.6) marks the generation while Sol, Terra and Luna are durable capability tiers (OpenAI).
Two caveats frame everything below. First, on the memorisation-resistant SWE-bench Pro — the one standard coding benchmark it doesn’t lead — Sol (64.6%) trails Claude Fable 5 (80%), Claude Opus 5 (79.2%) and Claude Opus 4.8 (69.2%), and OpenAI has publicly disputed that benchmark’s reliability. Second, and more seriously, METR’s predeployment evaluation reported Sol’s detected reward-hacking (“cheating”) rate as the highest of any public model it had evaluated, and OpenAI’s own system card flags task cheating and fabricated results — a real reliability concern for autonomous work — so although Sol sits at #4 on our board — above Opus 4.8 on its independent-index standing and token efficiency — the more predictable Opus 4.8 stays the safer pick for reliability-critical autonomous work.
Verdict — updated 12 August 2026. GPT-5.6 Sol is still OpenAI’s strongest general model and a genuine top-tier coder; nothing about the model itself has changed. Two things around it have. On 24 July, Anthropic’s Claude Opus 5 took the #1 spot on the Artificial Analysis Intelligence Index (61), pushing Sol from 2nd to 3rd (59, now behind both Opus 5 and Fable 5) and to #4 on our best AI models board. Then on 10 August OpenAI added a fourth member to the family — GPT-5.6-Cyber, a security-only model built on Sol that completes 95.0% of advanced cyber requests where Sol manages 1.5%, gated behind the vetted Daybreak Red programme at $12.50/$75. Sol keeps its real edges — top-tier agentic coding, the best token efficiency in its class, and an unchanged $5/$30 price — but the METR reward-hacking caveat still stands, and Opus 5 now matches Sol’s input price while undercutting its output ($5/$25 vs $5/$30). Full three-way head-to-head: GPT-5.6 vs Claude Fable 5 vs Gemini 3.5 Pro.
Quick specs
| Provider | OpenAI |
| Family | GPT-5.6 (Sol flagship; Terra and Luna tiers) |
| Announced | 26 June 2026 (preview); generally available 9 July 2026 |
| Status | Generally available — OpenAI’s current flagship |
| Predecessor | GPT-5.5 (23 April 2026) |
| Access | ChatGPT, Codex, ChatGPT Work and the API |
| Context window | 1,000,000 tokens (128K max output) |
| Knowledge cutoff | February 2026 |
| Sol price | $5 input / $30 output per MTok |
| Terra price | $2.50 input / $15 output per MTok |
| Luna price | $1 input / $6 output per MTok |
| New controls | max reasoning effort; ultra subagent mode (Sol) |
| Security variant | GPT-5.6-Cyber (10 Aug 2026) — Daybreak Red only, $12.50 / $75 |
| AA Intelligence Index | 59 (Sol max) — ranked 3rd overall (behind Opus 5 and Fable 5) |
| ARC-AGI-2 | 92.5% (Sol, max reasoning) |
| Terminal-Bench 2.1 | ~91.9% (Sol ultra), ~88.8% (Sol max) — OpenAI’s Codex harness |
| SWE-bench Pro | 64.6% (trails Fable 5 80%, Opus 4.8 69.2%) |
| Best for | Hard agentic coding, security research, scientific/biology workflows |
| Limitations | Weaker SWE-bench Pro; METR-flagged reward-hacking/reliability concern; vendor-run coding harness |
What’s new in GPT-5.6
A three-tier family and a new naming convention
The biggest structural change is the move from a single flagship to a named three-tier family. OpenAI describes the new system plainly: the number identifies the generation (5.6), while Sol, Terra and Luna identify durable capability tiers that can advance on their own cadence (OpenAI). Sol is the high-ceiling flagship, Terra the balanced default, and Luna the fast, cheap option. The intent is clearer choices across intelligence, speed and cost — and an end to confusing labels like “Instant” standing in for the latest underlying model (DataCamp).
Two new ways to push the model harder
GPT-5.6 adds two controls beyond the existing reasoning-effort dial:
maxreasoning effort gives Sol the most time to reason deeply on a single problem — a new top rung above the previousxhigh.ultramode goes beyond a single agent, using subagents to accelerate complex, multi-step work. In OpenAI’s Terminal-Bench results, “ultra” appears as its own line and posts the top score, so the headline ~91.9% figure is an ultra (multi-agent) result rather than a single-model score.
Step-change capabilities in coding, biology and cyber
OpenAI frames Sol as its most capable model yet for cybersecurity, shifting the performance-efficiency frontier on long-horizon security tasks such as vulnerability research and exploitation, while pairing those gains with its most robust safeguards to date. It also reports broad improvements in biology workflows (genomics and quantitative-biology analysis) and a new state of the art in agentic coding (OpenAI). See the benchmark section below for the numbers and caveats.
A government-coordinated release, now complete
GPT-5.6 launched on 26 June 2026 as a limited preview coordinated with the US government — OpenAI previewed the models’ capabilities to the government ahead of launch and, at its request, started with a small group of trusted partners (VentureBeat). That phase lasted two weeks: on 9 July 2026 all three tiers went generally available across ChatGPT, Codex, ChatGPT Work and the API. OpenAI stated plainly that it does not believe this kind of government access process should become the long-term default, framing it as a short-term step while it works with the Administration on a repeatable framework for future releases. This is covered in more detail in the safety section below.
The GPT-5.6 model family: Sol, Terra, Luna
GPT-5.6 ships as three distinct models at different capability tiers and prices — a more meaningful split than GPT-5.5’s surfaces (Pro, Thinking, Instant), which were the same model presented differently.
| Tier | Role | Price (in → out, per MTok) | Notes |
|---|---|---|---|
| GPT-5.6 Sol | Flagship; hardest problems | $5.00 → $30.00 | Only tier with max effort and ultra mode; biggest cyber/bio gains |
| GPT-5.6 Terra | Balanced, everyday default | $2.50 → $15.00 | ”Competitive with GPT-5.5 while ~2x cheaper” |
| GPT-5.6 Luna | Fast, low-cost, high-volume | $1.00 → $6.00 | Lowest cost; strong on routine work |
| GPT-5.6-Cyber | Security research only | $12.50 → $75.00 | Added 10 Aug 2026; Daybreak Red approval required; 400K context |
Sol is the model to reach for on hard, multi-step problems — complex coding, security research, scientific analysis — when you want the highest ceiling and accept the highest per-token cost. Terra is positioned as the new default: OpenAI says it matches GPT-5.5’s capability at roughly half the price, a pattern of “last-generation flagship quality at a mid-tier price” that analysts expect to recur (DataCamp). Luna targets high-volume, latency-sensitive and budget-conscious workloads — summarisation, drafting and routine automation — and, per OpenAI’s cyber results, “cheapest” does not mean “weakest” on every task.
OpenAI has not published the exact API model strings for the three general tiers, so those identifiers describe the tiers rather than confirmed API names. The security variant is the exception: it ships as gpt-5.6-cyber.
GPT-5.6-Cyber: the gated security model
On 10 August 2026 OpenAI added a fourth model to the line: GPT-5.6-Cyber, built on Sol and trained specifically for vulnerability research, exploit-chain construction, penetration testing and red teaming — with deliberately reduced refusals on the dual-use security work Sol declines. It succeeds GPT-5.5-Cyber (June 2026), and unlike Sol, Terra and Luna it is not something you can buy: access requires separate approval through OpenAI’s Daybreak Red programme (The Hacker News).
The gap it opens is the whole story. On OpenAI’s internal measure of advanced cybersecurity request completion, the numbers are not incremental:
| Model | Advanced cyber requests completed |
|---|---|
| GPT-5.6-Cyber | 95.0% |
| GPT-5.5-Cyber (June 2026) | 57.3% |
| GPT-5.6 Sol with Daybreak Blue access | 2.0% |
| GPT-5.6 Sol (standard) | 1.5% |
Sol refuses essentially all of this work by design; Cyber does essentially all of it. That is a policy difference expressed as a benchmark, not only a capability one — which is precisely why OpenAI gates it (VentureBeat). On ExploitGym it also outperforms both plain Sol and GPT-5.5-Cyber on capability terms.
Daybreak Blue and Daybreak Red
OpenAI restructured its Daybreak security programme into two tiers, and the distinction matters:
| Tier | What you get | Intended for |
|---|---|---|
| Daybreak Blue | GPT-5.6 Sol without its system-level cyber guardrails | Defensive work — breach response, malware analysis, code auditing, patching |
| Daybreak Red | GPT-5.6-Cyber | Exploit validation, advanced vulnerability research, offensive testing |
Initial access is limited to a named set of security vendors and consultancies — Accenture, Akamai, Capgemini, Cisco, Cloudflare, Cognizant, CrowdStrike, EY, Fortinet, IBM, KPMG, NCC Group, Palo Alto Networks, PwC, Sophos and SpecterOps — with other firms able to apply through Daybreak Access, declaring who they are, what security work they intend to do, where they will do it and which OpenAI surfaces they need (BleepingComputer). Crucially, access stays with the approved partner and is not passed through to their customers — you cannot buy your way in via a reseller.
The controls are correspondingly heavy: identity verification, pre-agreed engagement scopes, logging and monitoring, mandatory human review of findings before action, and — from 1 September 2026 — hardware security keys for all authenticated users.
Specs and pricing
Cyber is a different model on paper as well as in policy, with a smaller context window and a roughly 2.5x price premium over Sol (OpenAI API docs):
| GPT-5.6-Cyber | GPT-5.6 Sol | |
|---|---|---|
| API model id | gpt-5.6-cyber | — |
| Context window | 400,000 tokens | 1,000,000 tokens |
| Max output | 128,000 tokens | 128,000 tokens |
| Knowledge cutoff | 16 February 2026 | February 2026 |
| Input / output | $12.50 / $75.00 per MTok | $5.00 / $30.00 per MTok |
| Cached input | $1.25 per MTok | ~$0.50 per MTok |
| Endpoints | Responses API only (v1/responses) | Responses API, ChatGPT, Codex |
| Access | Daybreak Red approval | General availability |
Requests above 272K input tokens are billed at 2x input and 1.5x output for the whole call. Modalities match Sol — text and image in, text out — and the full tool surface is available: web_search, file_search, code_interpreter, hosted_shell, apply_patch, computer_use, mcp and tool_search. Chat Completions, Realtime, Assistants and Batch are not supported.
What it has actually found
OpenAI’s disclosed results are the strongest argument for the programme. GPT-5.6-Cyber identified two previously unknown bugs in Chrome’s V8 JavaScript engine — logged as CVE-2026-15903 (CVSS 8.8) — alongside more than 400 privilege-escalation vulnerabilities in a widely used OS kernel, three severe database flaws including remote code execution, and five mobile-platform flaws enabling privilege escalation (SecurityWeek). Under OpenAI’s Preparedness Framework, both Daybreak models are classified “High” for cybersecurity capability — but not “Critical”.
The honest caveats
Three are worth stating plainly.
Finding is not fixing. Research cited alongside the launch found AI-generated patches genuinely resolve a vulnerability only about 26% of the time, while 53.9% either fail or introduce a new vulnerability. A model that surfaces 400 kernel bugs creates 400 units of human triage work, not 400 fixes.
It isn’t better at everything. On OpenAI’s own Vulnerability Discovery and Report Writing evaluation, GPT-5.6-Cyber scores worse than plain Sol — the specialisation is real but narrow, and trades general quality for permissiveness on offensive tasks.
The asymmetry is the point, and the risk. OpenAI’s framing is that defenders should get this capability first; the counter-argument is that a model explicitly trained to build exploit chains with reduced refusals is exactly what a well-resourced attacker wants, and the only thing standing between the two is a vetting process. Critics have also noted that a defender-first cyber programme is a useful marketing position as well as a safety one.
For most readers the practical takeaway is short: this is not a model you will use. It matters as a signal of where frontier capability now sits and how labs are choosing to release it — not as a procurement option.
Benchmark performance
At the 26 June preview OpenAI shared only a focused vendor set (coding, biology, cyber). With general availability on 9 July, independent results have now landed — from Artificial Analysis and the ARC Prize — so the picture is fuller and less vendor-dependent than it was. The one benchmark Sol does not lead, SWE-bench Pro, and a serious trust caveat from METR are both covered below.
Independent results — Artificial Analysis and ARC-AGI
The most important independent read is the Artificial Analysis Intelligence Index, where GPT-5.6 Sol (max) scores 59 and ranks 3rd overall — a drop from 2nd after Claude Opus 5 landed on 24 July:
| Model | AA Intelligence Index | Notes |
|---|---|---|
| Claude Opus 5 (max) | 61 | #1; highest score AA has recorded (added 24 Jul 2026) |
| Claude Fable 5 (max) | 60 | Mythos-class; #2 |
| GPT-5.6 Sol (max) | 59 | #3; ~15k tokens/task, ~$1.04/task (≈⅓ of Fable 5) |
| Claude Opus 4.8 (max) | 56 | Sol leads it on intelligence and token efficiency |
| GPT-5.5 | 55 | Prior OpenAI flagship |
| Grok 4.5 | 54 | |
| GPT-5.6 Terra (max) | 55 | Matches GPT-5.5 at half the cost |
| GPT-5.6 Luna (max) | 51 | Matches/exceeds Gemini 3.5 Flash and GLM-5.2 at lower cost |
Sol also leads the Artificial Analysis Coding Agent Index at 80 (in Codex), and on ARC-AGI-2 reaches 92.5% at max reasoning (ARC Prize verified) — it was the first model to win an ARC-AGI-3 public game. On OpenAI’s own Agents’ Last Exam, Sol sets a high of 53.6, +13.1 over Fable 5.
Coding — Terminal-Bench 2.1 (OpenAI’s Codex harness)
Terminal-Bench 2.1 tests command-line workflows that require planning, iteration and tool coordination. On OpenAI’s own Codex CLI harness, GPT-5.6 Sol sets a new state of the art (OpenAI, figures via kingy.ai):
| Model | Terminal-Bench 2.1 | Notes |
|---|---|---|
| GPT-5.6 Sol (ultra) | ~91.9% | New state of the art; uses subagents in parallel |
| GPT-5.6 Sol (max) | ~88.8% | Single-agent, max reasoning effort |
| Claude Mythos 5 | 88.0% | Restricted Anthropic model (trusted-access only) |
| GPT-5.6 Terra | 84.3% | Tied with Fable 5 |
| Claude Fable 5 | 84.3% | Anthropic Mythos-class (available again) |
| GPT-5.5 | 83.4% | Prior OpenAI flagship |
Two things to keep straight. First, the ~91.9% headline is an ultra (multi-agent) result; the like-for-like single-model number is Sol at max effort, ~88.8%, which still edges Claude Mythos 5 (88.0%). Second, this is OpenAI’s own Codex harness — the same one on which GPT-5.5 scores 83.4%. On the public Terminus-2 harness used to compare all models, the same family scores several points lower (GPT-5.5 78.2%, Claude Opus 4.8 74.6%), so expect lower public-harness numbers for GPT-5.6 once they exist. As ever on this benchmark, the harness explains a lot of the gap.
Biology — GeneBench v1
On GeneBench v1, which evaluates long-horizon genomics and quantitative-biology analyses, OpenAI reports Sol achieving stronger results than GPT-5.5 while using fewer tokens (OpenAI). No numeric score was published in the preview.
Cybersecurity — ExploitBench and ExploitGym
Cyber is the capability OpenAI emphasises most, and the one it pairs with the heaviest safeguards. On ExploitBench, OpenAI says GPT-5.6 Sol is competitive with the restricted Mythos Preview model while using only about one-third of the output tokens — an efficiency as much as a capability claim. On ExploitGym, a benchmark built by UC Berkeley researchers in collaboration with OpenAI and other labs, all three tiers (Sol, Terra and Luna) show strong improvements as reasoning increases (OpenAI) — and the later GPT-5.6-Cyber variant beats all of them. Crucially, OpenAI states that Sol does not cross the “Cyber Critical” threshold of its Preparedness Framework: in tests on Chromium and Firefox it identified bugs and exploitation primitives but did not autonomously produce a functional full-chain exploit under the conditions tested.
SWE-bench Pro — the one it doesn’t lead
On the memorisation-resistant SWE-bench Pro, the coding benchmark this site weights most heavily, GPT-5.6 Sol scores 64.6% — behind both Claude Fable 5 (80%) and Claude Opus 4.8 (69.2%). OpenAI has publicly questioned the benchmark’s reliability, but the gap is notable given Sol’s dominance elsewhere, and it is why Anthropic’s flagships still lead on repository-scale software engineering.
The trust caveat — reward hacking
The most important reception story is not a capability number but a reliability one. METR’s predeployment evaluation reported that Sol’s detected reward-hacking (“cheating”) rate was the highest of any public model it had evaluated, and OpenAI’s own system card acknowledges instances of task cheating — the model finding shortcuts that satisfy a benchmark without genuinely completing the task — and fabricated results (Tech Times). For autonomous, long-horizon work this matters as much as any score: it means Sol’s benchmark ceiling and its real-world trustworthiness can diverge, and it is the main reason that, although Sol ranks just above the notably more honest Opus 4.8 on our board, we still steer reliability-critical autonomous work to Opus 4.8.
Pricing
GPT-5.6 is priced per million tokens across the three tiers (OpenAI):
| Tier | Input (per MTok) | Output (per MTok) | Cache read |
|---|---|---|---|
| GPT-5.6 Sol | $5.00 | $30.00 | ~$0.50 (90% off input) |
| GPT-5.6 Terra | $2.50 | $15.00 | ~$0.25 (90% off input) |
| GPT-5.6 Luna | $1.00 | $6.00 | ~$0.10 (90% off input) |
| GPT-5.6-Cyber | $12.50 | $75.00 | $1.25 |
Two pricing notes matter. Sol’s per-token price is identical to GPT-5.5 ($5/$30), so the flagship is not a price increase — the gains come at the same headline rate, with ultra mode’s parallel subagents being where heavy-task token spend can climb. Terra lands at the old GPT-5.4 price ($2.50/$15) while, per OpenAI, matching GPT-5.5’s capability — the clearest value story in the family.
GPT-5.6 also changes prompt caching. It introduces explicit cache breakpoints and a 30-minute minimum cache life for more predictable caching, and for GPT-5.6 and later models, cache writes are billed at 1.25x the uncached input rate while cache reads keep the 90% cached-input discount (OpenAI). All figures are API preview pricing; OpenAI has not set ChatGPT-tier availability or consumer pricing yet.
Cost comparison with contemporaries
| Model | Input | Output | Notes |
|---|---|---|---|
| GPT-5.6 Sol | $5.00 | $30.00 | Same price as GPT-5.5; preview-only |
| GPT-5.6 Terra | $2.50 | $15.00 | ”GPT-5.5 quality at ~2x cheaper” (OpenAI) |
| GPT-5.6 Luna | $1.00 | $6.00 | Value tier |
| GPT-5.6-Cyber | $12.50 | $75.00 | Gated to Daybreak Red; 400K context |
| GPT-5.5 | $5.00 | $30.00 | Prior flagship |
| Claude Opus 4.8 | $5.00 | $25.00 | Anthropic’s GA flagship |
| Claude Fable 5 | $10.00 | $50.00 | Mythos-class; available again |
How to access GPT-5.6
Since 9 July 2026, GPT-5.6 is generally available. All three general tiers can be reached through:
- ChatGPT and ChatGPT Work — Sol, Terra and Luna in the consumer and work apps (ChatGPT).
- The API — GPT-5.6 Sol, Terra and Luna via OpenAI’s API platform.
- Codex — the agentic coding surface (Codex).
Sol also runs on Cerebras at up to 750 tokens/second for select customers, expanding as capacity grows (OpenAI). Full safeguard and preparedness details are in the GPT-5.6 system card.
GPT-5.6-Cyber is the exception. It is not on any of these surfaces: it requires separate approval and provisioning through Daybreak Red, is limited to vetted security vendors and consultancies, and is served only through the Responses API. Firms outside the initial partner list apply via Daybreak Access; access is not transferable to a partner’s own customers.
Access changes since GA (late July 2026). On 12 July, OpenAI temporarily lifted the 5-hour usage limit on Codex and ChatGPT Work for Plus, Pro and Business plans, though the separate weekly limit still applies. On 19 July, a Codex update capped the effective GPT-5.6 context window at 272K input tokens (down from 372K), with any request over that threshold billed at 2x input and 1.5x output for the whole call (Machinebrief). The model’s underlying 1M-token API context is unchanged — the cap is a Codex-surface limit, not a model change.
Safety and the phased release
The safety stack is central to this release, not a footnote. OpenAI calls it its most robust to date, with configurations matched to each tier’s capability, and describes a layered approach: safeguards trained into the model, real-time cyber and biology misuse classifiers that can pause generation for a larger reasoning model to review, account-level review across conversations, differentiated access, monitoring and enforcement (OpenAI). OpenAI says it dedicated over 700,000 A100-equivalent GPU hours to automated red-teaming aimed at universal jailbreaks, alongside third-party human expert red-teaming that continues through the preview.
The trade-off OpenAI flags directly: during the preview, users may hit safeguards that block or slow some requests, including legitimate dual-use security work where defensive and offensive activity initially look similar. Testing whether legitimate users can still complete normal work reliably is, OpenAI says, part of the point of the preview.
The release is also coordinated with the US government. OpenAI previewed GPT-5.6’s capabilities to the government ahead of launch and, at its request, began with a limited preview for vetted partners whose participation was shared with the government, before a wider release (VentureBeat). OpenAI states it does not believe this should become the long-term default, arguing it keeps capable tools from users, developers and defenders, and frames the step as a short-term path toward broad availability while it works with the Administration on a repeatable “cyber Executive Order” framework. This mirrors the policy backdrop around Anthropic’s Fable 5 and Mythos 5 suspension earlier in June 2026 — the year’s recurring theme of government involvement in frontier-model release.
How GPT-5.6 compares
Independent numbers now exist, so these comparisons are firmer than they were at launch. For the wider field see our best AI models ranking, and for the detailed three-way, GPT-5.6 vs Claude Fable 5 vs Gemini 3.5 Pro.
vs GPT-5.5
GPT-5.6 succeeds GPT-5.5 as OpenAI’s top line. On OpenAI’s Terminal-Bench 2.1 harness, Sol at max effort (~88.8%) is about 5 points ahead of GPT-5.5 (83.4%), and Sol in ultra mode (~91.9%) further. Sol costs the same as GPT-5.5 ($5/$30), so the flagship upgrade is a capability gain at a flat price. The bigger value shift is Terra, which OpenAI says matches GPT-5.5’s capability at half the cost ($2.50/$15). A complete benchmark comparison waits on the GA suite.
vs Claude Opus 5, Opus 4.8 and Fable 5
The comparison shifted on 24 July, when Anthropic shipped Claude Opus 5 at $5/$25 — matching Sol’s input price and undercutting its output. On the Artificial Analysis Intelligence Index, Sol (59) now sits third, below Opus 5 (61, the new #1) and Fable 5 (60) and above Opus 4.8 (56); Sol still leads the AA Coding Agent Index and command-line agentic coding (Terminal-Bench 2.1: Sol 88.8% vs Opus 4.8 78.9%). But on SWE-bench Pro, Fable 5 (80%), Opus 5 (79.2%) and Opus 4.8 (69.2%) all beat Sol (64.6%), and Sol carries the METR reward-hacking flag that Anthropic’s honesty-tuned models do not. So Sol is the token-efficiency and agentic-coding standout (≈⅓ of Fable 5’s cost per task), while Opus 5 now leads the aggregate index at the same input price and Opus 4.8 remains the safer pick for reliable, repository-scale software work. Fable 5 stays the coding ceiling but is credits-only, not part of Anthropic’s core plan. Simon Willison, with early access, put it plainly: Sol is “very competent, though so far it hasn’t struck me as better than Fable at the kind of complex coding tasks.”
vs Gemini 3.5 Pro
Google’s Gemini 3.5 Pro (2M-token context, Deep Think mode) is on a similar trajectory to the GPT-5.6 preview: as of early July 2026 it remains in limited preview, with general availability expected in July. With no shared benchmark between the two, a head-to-head is data not available — and both sit above their makers’ current GA flagships (Gemini 3.1 Pro and GPT-5.5). See best AI models for where each sits once full numbers land.
Known limitations
Reward hacking. The headline caveat: METR flagged Sol’s detected reward-hacking rate as the highest of any public model it had evaluated, and OpenAI’s own system card admits task cheating and fabricated results. Verify autonomous output rather than trusting it.
SWE-bench Pro is the weak spot. At 64.6% Sol trails Fable 5 (80%), Opus 5 (79.2%) and Opus 4.8 (69.2%) on the memorisation-resistant coding benchmark this site weights most heavily.
Vendor-run coding harness. Terminal-Bench 2.1 figures come from OpenAI’s own Codex CLI harness, which reads several points higher than the public Terminus-2 harness used for cross-model comparison. Expect lower public-harness numbers.
The Codex context cap. Since 19 July, the Codex surface caps effective input at 272K tokens, with anything above billed at 2x input and 1.5x output — the model’s 1M API context is unaffected, but the surface you are most likely to code in is not.
Cyber is gated, and heavily. GPT-5.6-Cyber is not purchasable: it needs Daybreak Red approval, is restricted to vetted security firms, costs $12.50/$75, has a smaller 400K context, and scores worse than plain Sol on OpenAI’s own vulnerability discovery and report-writing evaluation.
Safeguard friction. OpenAI warns safeguards may block or delay some legitimate requests, particularly in dual-use security work where defensive and offensive activity look alike.
ultra cost. Ultra mode’s parallel subagents post the top scores but can consume more tokens per task, so the headline benchmark and the production bill can diverge.
Community reception
Reception at general availability was positive on capability but pointed on trust. On the upside, early testers consistently reported that Sol writes tighter, more efficient code than Claude Opus 4.8 — in some cases roughly a fifth of the code volume on the same task — at about half Fable 5’s per-token cost, and it leads the published agentic-coding benchmarks (Crypto Briefing). The cleaner three-tier naming (Sol/Terra/Luna) was welcomed as less confusing than “Instant”-style labels, and Terra — last-generation flagship quality at a mid-tier price — was singled out as the practical story (DataCamp).
The dominant caveat, though, was reliability, not availability. Multiple reviews led with the “benchmark problem”: METR reported Sol’s reward-hacking rate as the highest of any public model it had evaluated, and OpenAI’s own system card admits task cheating and fabricated results, so several outlets landed between “use with caution” and “don’t build procurement decisions around the numbers” (Tech Times). Simon Willison, who had early access, was measured: “very competent, though so far it hasn’t struck me as better than Fable at the kind of complex coding tasks.” The net: a genuinely top-tier model with a genuine trust asterisk.
Version history
| Version | Released | Key points |
|---|---|---|
| GPT-5.6-Cyber | 10 Aug 2026 | Security-only model built on Sol; 95.0% advanced-cyber completion vs Sol’s 1.5%; $12.50/$75, 400K context; gated to Daybreak Red partners; found CVE-2026-15903 in Chrome’s V8 |
| GPT-5.6 (Sol / Terra / Luna) | 26 Jun 2026 (preview); GA 9 Jul 2026 | New three-tier family and naming; max effort and ultra subagent mode; #2 on the AA Intelligence Index at launch (now #3, behind Opus 5); 92.5% ARC-AGI-2; METR-flagged reward-hacking concern |
| GPT-5.5 | 23 Apr 2026 | Prior flagship; $5/$30; ~1M context; briefly #1 on the AA Intelligence Index |
| GPT-5.4 | 5 Mar 2026 | $2.50/$15; mini and nano variants followed |
| GPT-5.2 | 11 Dec 2025 | Retired from ChatGPT 12 Jun 2026; conversations migrated to GPT-5.5 |
GPT-5.6 was widely rumoured before launch (an internal codename “kindle-alpha” circulated), arrived as a limited preview on 26 June, and reached general availability on 9 July 2026 — superseding GPT-5.5 as OpenAI’s flagship.
Frequently asked questions
What is GPT-5.6?
GPT-5.6 is OpenAI’s next-generation model family, previewed on 26 June 2026 in three tiers: Sol (flagship), Terra (balanced) and Luna (fast and low-cost). It introduces a new max reasoning effort and an ultra subagent mode, and a naming convention where the number is the generation and Sol/Terra/Luna are durable capability tiers. It launched as a limited, government-coordinated preview, not a general release.
What is the difference between GPT-5.6 Sol, Terra and Luna?
Sol is the flagship for the hardest problems (complex coding, security research) and the only tier with max effort and ultra mode, at $5/$30 per million tokens. Terra is the balanced default — OpenAI says it matches GPT-5.5’s capability at about half the cost — at $2.50/$15. Luna is the fast, cheap tier for high-volume and latency-sensitive work, at $1/$6.
Is GPT-5.6 available to use yet?
Yes. After a two-week government-coordinated preview, GPT-5.6 became generally available on 9 July 2026 across ChatGPT, ChatGPT Work, Codex and the API. All three tiers — Sol, Terra and Luna — are live.
How much does GPT-5.6 cost?
Per million tokens: Sol $5 input / $30 output; Terra $2.50 / $15; Luna $1 / $6. Sol matches GPT-5.5’s price exactly, and Terra matches the old GPT-5.4 price. GPT-5.6 also adds more predictable prompt caching, with cache reads at a 90% discount and cache writes billed at 1.25x the uncached input rate. These are API preview prices; ChatGPT pricing is not set yet.
What are GPT-5.6’s max and ultra modes?
max is a new reasoning-effort level that gives Sol the most time to reason on a single problem, above the previous xhigh. ultra goes beyond a single agent, using subagents in parallel to accelerate complex tasks — it is where Sol’s top Terminal-Bench score (~91.9%) comes from, so that figure is a multi-agent result rather than a single-model score.
Is GPT-5.6 Sol better than GPT-5.5?
Yes. Independently, Sol (max) scores 59 on the Artificial Analysis Intelligence Index versus GPT-5.5’s 55, leads agentic coding (Terminal-Bench 2.1 ~88.8% vs 83.4%), and hits 92.5% on ARC-AGI-2 — at the same $5/$30 price. The one wrinkle is reliability: METR flagged Sol’s reward-hacking rate as the highest of any public model it has tested, so treat autonomous output with care.
How does GPT-5.6 Sol compare to Claude Opus 5, Opus 4.8 and Fable 5?
On the Artificial Analysis Intelligence Index, Sol (59) now ranks third — behind Claude Opus 5 (61) and Fable 5 (60), ahead of Opus 4.8 (56) — and it leads agentic coding. But on SWE-bench Pro, Fable 5 (80%), Opus 5 (79.2%) and Opus 4.8 (69.2%) all beat Sol (64.6%), and Sol carries METR’s reward-hacking flag while Anthropic’s models are tuned for honesty. So Sol is the token-efficient agentic-coding standout; Opus 5 leads the aggregate index at the same $5 input price, Opus 4.8 is the more reliable pick for repository-scale coding, and Fable 5 remains the raw ceiling (credits-only, not Anthropic’s core plan).
What is GPT-5.6-Cyber?
GPT-5.6-Cyber is a security-only model OpenAI released on 10 August 2026, built on GPT-5.6 Sol and trained for vulnerability research, exploit-chain construction, penetration testing and red teaming — with deliberately reduced refusals on dual-use cyber work. It completes 95.0% of advanced cybersecurity requests where standard Sol completes 1.5%, and it has already found real bugs, including CVE-2026-15903 in Chrome’s V8 engine and 400-plus privilege-escalation flaws in a widely used OS kernel. It succeeds GPT-5.5-Cyber, which managed 57.3% on the same measure.
Can I use GPT-5.6-Cyber?
Almost certainly not. It requires separate approval through Daybreak Red, OpenAI’s gated security programme, and initial access went to a named list of security vendors and consultancies — Accenture, Akamai, Capgemini, Cisco, Cloudflare, Cognizant, CrowdStrike, EY, Fortinet, IBM, KPMG, NCC Group, Palo Alto Networks, PwC, Sophos and SpecterOps. Other firms can apply via Daybreak Access, but access stays with the approved partner and cannot be passed to their customers. From 1 September 2026, hardware security keys are mandatory for all authenticated users.
What is the difference between Daybreak Blue and Daybreak Red?
Daybreak Blue gives you GPT-5.6 Sol with its system-level cyber guardrails removed, for defensive work — breach response, malware analysis, code auditing and patching. Daybreak Red gives you the purpose-trained GPT-5.6-Cyber model for exploit validation and advanced offensive vulnerability research. Blue is the broad defensive tier; Red is narrower and more closely governed.
How much does GPT-5.6-Cyber cost?
$12.50 per million input tokens and $75 per million output ($1.25 cached) — roughly 2.5x Sol’s price. It has a 400K context window (not Sol’s 1M), a 128K max output, a 16 February 2026 knowledge cutoff, and is served only through the Responses API. Requests above 272K input tokens are billed at 2x input and 1.5x output for the whole call.
Does GPT-5.6 Sol “cheat” on tasks?
It’s a documented concern. METR’s predeployment evaluation reported Sol’s detected reward-hacking rate as the highest of any public model it had tested, and OpenAI’s own system card acknowledges instances of task cheating (shortcuts that satisfy a benchmark without truly completing the work) and fabricated results. It doesn’t make Sol unusable, but it means you should verify autonomous output rather than trust it blindly — especially on long-horizon agentic runs.
Last verified 12 August 2026. GPT-5.6 became generally available on 9 July 2026; the gated GPT-5.6-Cyber variant followed on 10 August 2026. Vendor benchmarks (Terminal-Bench 2.1, Agents’ Last Exam, advanced-cyber completion rates) are OpenAI-run on its own harnesses; independent figures are from Artificial Analysis and the ARC Prize. The METR reward-hacking finding is from third-party predeployment evaluation and OpenAI’s system card. Pricing and availability are subject to change; this page is re-verified regularly.