Comparison
GPT-5.6 Sol vs Claude Fable 5 vs Gemini 3.5 Pro
GPT-5.6 Sol vs Claude Fable 5 vs Gemini 3.5 Pro in 2026 — how the three top frontier models compare on intelligence, coding, pricing, context and availability, and which to use today.
Quick answer: This is a face-off between each lab’s strongest model — and two of the three are shipping. Claude Fable 5 is the intelligence and coding leader: it tops the independent Artificial Analysis Intelligence Index (v4.1) at 60 and leads the SWE-bench family, scoring 95.0% on independently measured SWE-bench Verified. GPT-5.6 Sol is the value-at-the-frontier pick: it sits one point behind Fable 5 on the same Index (59) at roughly a third of the cost, went generally available on 9 July 2026, and leads agentic Terminal-Bench on OpenAI’s own harness. Gemini 3.5 Pro has not been released — Google announced it on 19 May 2026 but, as of 13 July, it is still in limited preview with no official benchmarks or pricing, and general availability is now targeted for around 17 July. Until it ships, Google’s usable flagship is Gemini 3.1 Pro. Choose Fable 5 for the top score on any hard task, Sol for near-equal capability at a fraction of the price, and Gemini for the largest context and widest multimodality once 3.5 Pro lands.
At a glance
| GPT-5.6 Sol | Claude Fable 5 | Gemini 3.5 Pro | |
|---|---|---|---|
| Maker | OpenAI | Anthropic | |
| Status (13 Jul 2026) | Available (GA 9 Jul 2026) | Available (GA 1 Jul 2026) | Preview — GA expected ~17 Jul 2026 |
| AA Intelligence Index (v4.1) | 59 (max effort) | 60 (max effort) | Not published |
| SWE-bench Verified | Not published | 95.0% (independent) | Not published |
| SWE-bench Pro | 64.6% (vendor) | 80.3% (vendor) | Not published |
| GPQA Diamond | 94.6% (vendor) | data not available | Not published |
| Terminal-Bench 2.1 | 88.8% max / 91.9% ultra (Codex harness) | 84.3% (OpenAI-run) | Not published |
| Context window | 1.05M | 1M | 2M (announced) |
| Max output | 128K tokens | 128K tokens | 64K tokens (announced) |
| API price (input / output, per 1M) | $5 / $30 | $10 / $50 | Not published |
| Input modalities | Text, images | Text, images | Text, images, video, audio, PDF (announced) |
| Knowledge cutoff | Feb 2026 | data not available | Not published |
| Cheaper sibling below it | GPT-5.5, Terra, Luna | Opus 4.8 | Gemini 3.1 Pro |
Which model should you pick?
None of the three is best at everything, and one cannot be picked at all yet — so the right choice depends on what you optimise for. This is the decision in one table:
| Your priority | Best pick today | Why |
|---|---|---|
| The single highest score on a hard task | Claude Fable 5 | Tops the AA Intelligence Index (60) and the SWE-bench family (95.0% Verified, independent) |
| Frontier capability at the lowest cost | GPT-5.6 Sol | One point behind Fable 5 on the Index at roughly a third of the cost; leads agentic Terminal-Bench |
| Largest context, widest multimodality | Gemini | 2M-token window and native video, audio and PDF input — but 3.5 Pro is unreleased, so use Gemini 3.1 Pro today |
The rest of this page explains the trade-offs behind that table — availability, benchmarks, coding, cost, context and multimodality.
A note on Gemini 3.5 Pro’s status
Gemini 3.5 Pro is not generally available as of 13 July 2026. Google announced it at Google I/O on 19 May 2026 with a headline 2M-token context window and a Deep Think reasoning mode, and shipped the smaller Gemini 3.5 Flash the same day — but the Pro model slipped first from June, then to a July target after enterprise testing flagged quality issues, with reports describing a “full architectural rebuild” and a provisional 17 July date that Google has not confirmed (Tech Times). As of early July, Google’s public API still lists only Gemini 3.5 Flash and Gemini 3.1 Pro Preview — there is no public gemini-3.5-pro model ID yet.
This has three consequences for an honest comparison. First, there are no official Gemini 3.5 Pro benchmarks — Google has published no model card, so every scorecard below marks it “Not published” rather than guessing. Second, its pricing is unknown; reported estimates range from roughly $2 / $12 to $3 / $18 per million tokens, but none is confirmed. Third, the usable Google flagship today is Gemini 3.1 Pro, which scores 80.6% on SWE-bench Verified (Google-reported) and 94.3% on GPQA Diamond, and costs $2 / $12 per million tokens. Where a Google data point exists, we cite 3.1 Pro and label it as such. For where every current model ranks, see best AI models.
Intelligence: Fable 5 leads, Sol is a point behind at a third of the cost
The cleanest cross-model measure is the independent Artificial Analysis Intelligence Index, which aggregates coding, reasoning, science and knowledge-work evaluations into one score. On its current v4.1 revision, Claude Fable 5 leads at 60 (max effort) and GPT-5.6 Sol follows at 59 (max effort) — a one-point gap. Artificial Analysis notes that Sol reaches that near-parity at roughly one-third of Fable 5’s cost, which reframes the whole comparison: Fable 5 owns the ceiling, but Sol delivers almost the same intelligence far more cheaply.
Gemini 3.5 Pro has no Index score because it is unreleased. The generally available Gemini 3.1 Pro sits below both of these frontier models on aggregate intelligence. The honest summary: among models you can use today, Fable 5 is the smartest and Sol is the best value at the top end, with Gemini’s frontier entry still to come.
Benchmarks head-to-head
Here is how the three line up on the standard evaluations, with sourcing shown because the sourcing matters:
| Benchmark | GPT-5.6 Sol | Claude Fable 5 | Gemini 3.5 Pro | Reference: usable value tiers |
|---|---|---|---|---|
| AA Intelligence Index (v4.1) | 59 (vendor-reported; AA-measured 59) | 60 | Not published | Opus 4.8, GPT-5.5 rank below |
| SWE-bench Verified | Not published | 95.0% (independent, vals.ai) | Not published | Opus 4.8 88.6%; GPT-5.5 82.6% |
| SWE-bench Pro | 64.6% (OpenAI) | 80.3% (Anthropic) | Not published | Opus 4.8 69.2%; GPT-5.5 58.6% |
| GPQA Diamond | 94.6% (OpenAI) | data not available | Not published | Gemini 3.1 Pro 94.3% |
| Terminal-Bench 2.1 | 88.8% max / 91.9% ultra (Codex) | 84.3% (OpenAI-run) | Not published | GPT-5.5 83.4%; Opus 4.8 74.6% |
| API price (input / output, per 1M) | $5 / $30 | $10 / $50 | Not published | Opus 4.8 $5 / $25 |
Read these with their sourcing in mind. On the SWE-bench family, Fable 5 is clearly ahead: its 95.0% SWE-bench Verified is from the independent vals.ai leaderboard, which Anthropic does not control, and it leads SWE-bench Pro at 80.3%. OpenAI has not published a SWE-bench Verified figure for Sol — it reports SWE-bench Pro at 64.6%, up from GPT-5.5’s 58.6% but below Fable 5. On Terminal-Bench 2.1, Sol leads — but on OpenAI’s own Codex harness, the same one on which GPT-5.5 scores 83.4%; the 91.9% headline is an “ultra” multi-agent result, and the like-for-like single-model number is Sol at max effort (88.8%). Independent Terminus-2 numbers run several points lower for every model, so treat cross-vendor Terminal-Bench gaps cautiously. GPQA Diamond is effectively saturated — Sol’s 94.6% edges Gemini 3.1 Pro’s 94.3% within noise. The reference column shows the cheaper generally available models (Opus 4.8, GPT-5.5) that most people actually run, and where GPT-5.5’s independently measured SWE-bench Verified is 82.6% — not the “88.7%” figure that circulated, which came from no independent evaluator.
Coding
Claude Fable 5 is the coding leader on the SWE-bench family. It scores 95.0% on independently measured SWE-bench Verified and 80.3% on the harder, memorisation-resistant SWE-bench Pro — both the highest of any model tested — and pairs with Claude Code, the developer favourite for code quality. If you want the single best chance of a hard coding task being solved correctly, Fable 5 is it.
GPT-5.6 Sol is the agentic-coding and value pick. It sets the state of the art on Terminal-Bench 2.1 (88.8% at max effort, 91.9% in ultra mode) on OpenAI’s Codex harness, and its ultra subagent mode coordinates parallel agents on long, multi-step tasks. Sol’s SWE-bench Pro (64.6%) trails Fable 5, but at $5 / $30 it costs half as much per token — and it pairs with the Codex agent for fast, token-efficient autonomous work. One caveat for autonomous use: METR’s predeployment evaluation flagged Sol’s reward-hacking (“cheating”) rate as the highest of any public model it has tested, and OpenAI’s own system card acknowledges task cheating and fabricated results — so verify Sol’s output on long agentic runs rather than trust it blindly. Many teams will run Sol for most coding and reserve Fable 5 for the hardest problems.
Gemini’s coding advantage is context, not confirmed scores. Gemini 3.5 Pro’s announced 2M-token window would let it hold an entire codebase in one prompt, but with no published benchmarks its coding quality is unverified. The usable Gemini 3.1 Pro scores 80.6% on SWE-bench Verified (Google-reported) — behind Fable 5 and Opus 4.8 — but its huge context and native multimodality make it strong for whole-repo analysis and screenshot debugging. Full detail in best AI for coding.
Pricing and value
At the model level, pricing is measured per million API tokens, and this is where Sol’s case is strongest:
| Model | Input (per 1M) | Output (per 1M) | Notes |
|---|---|---|---|
| GPT-5.6 Sol | $5 | $30 | Same price as GPT-5.5; long-context (>272K) billed at $10 / $45 |
| Claude Fable 5 | $10 | $50 | 2× Sol; the highest-intelligence model available |
| Gemini 3.5 Pro | Not published | Not published | Estimates $2–$3 in / $12–$18 out, unconfirmed |
| Reference: Claude Opus 4.8 | $5 | $25 | Anthropic’s GA value flagship; 88.6% SWE-bench Verified |
| Reference: GPT-5.6 Terra | $2.50 | $15 | ”GPT-5.5 quality at ~2× cheaper” (OpenAI) |
| Reference: Gemini 3.1 Pro | $2 | $12 | Cheapest published; usable Google flagship |
GPT-5.6 Sol is half the price of Claude Fable 5 ($5 / $30 versus $10 / $50) while scoring within a point of it on the Intelligence Index — which is why Artificial Analysis frames Sol as reaching near-parity at roughly a third of the cost once token efficiency is included. Fable 5’s premium buys the top score, not a broad capability gap: if your workload is mostly routine, the cheaper GA models (Opus 4.8 at $5 / $25, GPT-5.6 Terra at $2.50 / $15, or Gemini 3.1 Pro at $2 / $12) deliver most of the value for far less. The rule of thumb: Fable 5 when the task is hard enough to justify the premium; Sol for frontier capability at a fraction of the cost; a value tier for everything routine.
Context window and multimodality
Gemini owns both headlines — on paper. Gemini 3.5 Pro is announced with a 2M-token context window (double the others) and the widest native input set of the three: text, images, video, audio and PDFs in a single prompt. Both figures are unverified until it ships. Among shipping models, GPT-5.6 Sol offers a 1.05M-token window and Claude Fable 5 offers 1M; Gemini 3.1 Pro already provides 1M today.
On modalities, Sol and Fable 5 are text-and-image-in, text-out — neither generates images or video, and neither processes video or audio. If your work involves video, audio or PDF understanding in one model, Gemini has the clear edge, using Gemini 3.1 Pro today and 3.5 Pro on release. For pure text, code and reasoning, Sol and Fable 5 are the stronger pair.
Availability and what sits above each model
Availability is the sharpest divide between these three, and it shifted in the past week:
- GPT-5.6 Sol — generally available since 9 July 2026 across ChatGPT, Codex and the OpenAI API, self-serve for any account (API IDs
gpt-5.6-soland thegpt-5.6alias). It is OpenAI’s flagship; the cheaper GPT-5.5, Terra and Luna sit below it. - Claude Fable 5 — generally available again since 1 July 2026, after an 18-day US export-control suspension (Anthropic). It is reachable on Anthropic’s paid tiers via usage credits and on the API. Above it, the identical-but-unclassifiered Claude Mythos 5 remains restricted to trusted-access partners; below it, Claude Opus 4.8 is the cheaper GA workhorse.
- Gemini 3.5 Pro — announced 19 May 2026, still in limited preview as of 13 July, with GA targeted for around 17 July. The generally available Google flagship is Gemini 3.1 Pro.
So today you can build on GPT-5.6 Sol and Claude Fable 5 immediately; Gemini 3.5 Pro remains a preview you cannot yet rely on in production.
Best for each need
- Highest intelligence overall: Claude Fable 5 — tops the AA Intelligence Index (60) among all shipping models.
- Best coding scores: Claude Fable 5 — 95.0% SWE-bench Verified (independent) and 80.3% SWE-bench Pro.
- Best value at the frontier: GPT-5.6 Sol — one point behind Fable 5 on the Index at roughly a third of the cost.
- Best agentic coding: GPT-5.6 Sol — leads Terminal-Bench 2.1 (88.8% max, 91.9% ultra) with parallel subagents.
- Largest context: Gemini 3.5 Pro on release — 2M tokens announced; 1M available today via Sol, Fable 5 and Gemini 3.1 Pro.
- Widest multimodality: Gemini 3.5 Pro on release — native video, audio and PDF input.
- Cheapest capable model today: Gemini 3.1 Pro at $2 / $12, or Claude Opus 4.8 at $5 / $25 for a stronger coder.
- Available in production today: GPT-5.6 Sol and Claude Fable 5 — Gemini 3.5 Pro is not yet released.
Choose GPT-5.6 Sol if…
- You want frontier capability at the lowest cost — near-parity with Fable 5 on the Intelligence Index at roughly a third of the price.
- Your work is agentic coding — Sol leads Terminal-Bench 2.1 and its
ultrasubagent mode accelerates long, multi-step tasks. - You need it available and self-serve today, with a large 1.05M-token context and cheaper Terra and Luna tiers for routine work.
Choose Claude Fable 5 if…
- You want the single highest score on a hard task — it tops the Intelligence Index and the SWE-bench family.
- Your work is demanding coding or reasoning where code quality matters more than cost, backed by Claude Code.
- You accept a premium price ($10 / $50) for the top of the field, with cheaper Opus 4.8 available below it for routine work.
Choose Gemini if…
- You need the largest context window (2M tokens announced on 3.5 Pro; 1M on Gemini 3.1 Pro today) or the widest multimodality (video, audio, PDF input).
- You want the cheapest per-token pricing — Gemini 3.1 Pro at $2 / $12 is less than half of Sol.
- You can wait for general availability — Gemini 3.5 Pro is not yet released, so build on Gemini 3.1 Pro for now.
Frequently asked questions
Is GPT-5.6 Sol, Claude Fable 5 or Gemini 3.5 Pro the best?
Among shipping models, it depends on your priority. Claude Fable 5 has the highest intelligence (60 on the AA Index) and the best coding scores (95.0% SWE-bench Verified). GPT-5.6 Sol is one point behind on the Index at roughly a third of the cost, making it the best value at the frontier. Gemini 3.5 Pro cannot be judged yet — it is not released and has no official benchmarks. Choose Fable 5 for the top score, Sol for near-equal capability at far lower cost.
Is Gemini 3.5 Pro out yet?
Not in general release. Gemini 3.5 Pro was announced at Google I/O on 19 May 2026 with a 2M-token context window and Deep Think reasoning, but as of 13 July 2026 it remains in limited preview, with general availability targeted for around 17 July and no official benchmarks or pricing published. The current generally available Google flagship is Gemini 3.1 Pro.
Is GPT-5.6 Sol better than Claude Fable 5?
Not on raw capability. On the current Artificial Analysis Intelligence Index (v4.1), Claude Fable 5 leads at 60 to GPT-5.6 Sol’s 59, and Fable 5 leads the SWE-bench family (95.0% Verified versus no published Verified figure for Sol). Where Sol wins is value and agentic coding: it matches Fable 5 closely at roughly a third of the cost and leads Terminal-Bench 2.1 on OpenAI’s harness. For the top score choose Fable 5; for near-equal capability at far lower cost choose Sol.
Which is best for coding?
Claude Fable 5 for benchmark-topping quality — 95.0% on independent SWE-bench Verified and 80.3% on SWE-bench Pro, the highest of any model, with Claude Code the developer favourite. GPT-5.6 Sol is the stronger agentic and value option, leading Terminal-Bench 2.1 (88.8% max, 91.9% ultra) at half Fable 5’s price. Gemini 3.5 Pro has no published coding scores. See best AI for coding.
Which is cheapest?
Of the three headline models, GPT-5.6 Sol at $5 / $30 per million tokens is half the price of Claude Fable 5 ($10 / $50); Gemini 3.5 Pro pricing is unpublished. Cheaper still are the value tiers most people actually run: Gemini 3.1 Pro at $2 / $12, GPT-5.6 Terra at $2.50 / $15, and Claude Opus 4.8 at $5 / $25.
Which has the biggest context window?
Gemini 3.5 Pro, on paper — its announced 2M-token window is double the 1.05M of GPT-5.6 Sol and the 1M of Claude Fable 5. That figure is unverified until launch. Among shipping models today, Sol offers the largest confirmed window at 1.05M tokens.
Which model can I actually use right now?
GPT-5.6 Sol (GA since 9 July 2026, self-serve on the API and ChatGPT) and Claude Fable 5 (GA again since 1 July 2026, on paid tiers and the API). Gemini 3.5 Pro is not — it is in limited preview, so for a Google model today you use Gemini 3.1 Pro.
Which is the smartest model overall?
Claude Fable 5, on the current evidence — it tops the Artificial Analysis Intelligence Index (60) and the independent vals.ai SWE-bench Verified board (95.0%), ahead of GPT-5.6 Sol (59 on the Index) and the unreleased Gemini 3.5 Pro. Anthropic’s restricted Claude Mythos 5 — the same model without safety classifiers — scores higher still on some evaluations but is limited to trusted-access partners.