THE AI RANKINGS

xAI

Grok 4.6

Provider
xAI
Status
Current
Context
500,000 tok
Price
$2 / $6 /MTok

Grok 4.6 is xAI’s flagship model, launched on 12 August 2026 across Cursor, Grok Build and the SpaceXAI API, with partner availability on OpenRouter, Vercel and Cloudflare from day one. It succeeds Grok 4.5 after just five weeks, and it is built on the same reported ~1.5-trillion-parameter foundation — SpaceXAI attributes the gains almost entirely to post-training: curated model-generated reasoning data, regenerated supervised fine-tuning trajectories, and reinforcement learning on long agentic tasks spanning knowledge work, coding and domain-specific environments. The focus is long-running agents: SpaceXAI reports the model doing more self-testing and verification on extended trajectories, checking its own work before moving to the next step.

The launch claim is bigger than Grok 4.5’s — and this time the composite is independently confirmed. Artificial Analysis scores Grok 4.6 (high) at 61 on its Intelligence Index, tied with GPT-5.6 Sol Max for 3rd and one point behind Fable 5 Max (62), with Claude Opus 5 max on top (63) — the first time an xAI model has joined the frontier cluster, VentureBeat notes, overtaking Kimi K3. Its sharpest edge is economics: at $2/$6 per million tokens it posts the frontier’s lowest cost per task ($0.84 on Artificial Analysis’s index, using ~53 agent turns where Opus 5 uses ~103). The caveats are equally clear: it publishes no SWE-bench figure of any kind, and it loses to GPT-5.6 Sol Max on DeepSWE v1.1 and Terminal-Bench 3.0.

Quick specs

ProviderxAI / SpaceXAI
Released12 August 2026 (Cursor, Grok Build, API, OpenRouter, Vercel, Cloudflare)
API model IDgrok-4.6 (OpenAI-compatible endpoint)
Context window500,000 tokens (unchanged from Grok 4.5)
Parameters~1.5 trillion (reported; same foundation as Grok 4.5)
Knowledge cutoffNot disclosed
Input price$2.00 / MTok ($4.00 above 200K input)
Output price$6.00 / MTok ($12.00 above 200K input)
Cached input$0.50 / MTok
ModalitiesText and image in; text out
AA Intelligence Index61 (independent; tied 3rd with GPT-5.6 Sol Max)
GDPVal-AA v21753 Elo (2nd, behind only Claude Opus 5)
Cost per AA index task$0.84 — the lowest on the frontier
Best forCheap long-running agents, knowledge work, agentic coding at volume
LimitationsNo SWE-bench figure at all; weak Terminal-Bench 3.0 (26%); long-context calls double the price; EU availability unconfirmed at launch

TRY GROK 4.6 →

What’s new in Grok 4.6

Grok 4.6 is a post-training release, not a new base model. SpaceXAI says it kept Grok 4.5’s foundation and changed the recipe: curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, an improved optimiser, regenerated SFT trajectories, and agentic reinforcement learning across knowledge work, general coding and domain-specific environments.

Built for long-running agents

Where Grok 4.5 was pitched at coding, Grok 4.6 is pitched at long agentic tasks — researching a topic, working across an entire codebase, or turning an idea into a working application over many iterations. The emergent behaviour SpaceXAI highlights is self-verification: on longer trajectories the model reviews its own work before moving on, a behaviour it attributes to reinforcement learning on long tasks rather than explicit instruction. The improvement over its predecessor is across the board — Grok 4.6 beats Grok 4.5 on all nine benchmarks in SpaceXAI’s launch table, including a 227-point jump on GDPVal-AA v2 (1753 vs 1526) and a 10-point jump on APEX-Agents (57.5% vs 47.1%).

The frontier’s cheapest completed task

Grok 4.6 keeps Grok 4.5’s headline economics and extends them. Artificial Analysis measures $0.84 per Intelligence Index task — under Kimi K3 ($0.94), GPT-5.6 Sol ($1.04) and far under the Claude line — driven by turn efficiency: ~53 agent turns and ~0.5B input tokens on average, versus ~103 turns and ~2.0B input tokens for Claude Opus 5. For long agentic runs, where input tokens compound every turn, that efficiency matters more than the sticker price.

Same window, same modalities

The trade-offs Grok 4.5 made are unchanged: a 500K context window (still half Grok 4.3’s 1M), text and image input only (no video), and text output. Reasoning effort is configurable, with “high” the setting behind the published benchmarks. New this release is a Grok 4.6 Fast variant at double the standard price for latency-sensitive work.

Benchmark performance

Two things distinguish this launch from Grok 4.5’s. First, the composite is independently verified on day one: Artificial Analysis’s own run puts Grok 4.6 at 61, exactly matching SpaceXAI’s table. Second, the vendor table’s competitor numbers are pulled from published system cards and leaderboards, not SpaceXAI reruns — more comparable than a tuned in-house harness, but still SpaceXAI’s selection of which benchmarks to show.

SpaceXAI’s launch table (12 August 2026)

BenchmarkGrok 4.6 (high)Grok 4.5 (high)GPT-5.6 Sol MaxFable 5 Max
AA Intelligence Index61566162
GDPVal-AA v21753152617281741
CursorBench v3.269.9%66.7%67.2%70.5%
DeepSWE v1.165.9%54%73%70%
FrontierCode v1.161.3%56.6%60.6%63.6%
APEX-Agents57.5%47.1%56.7%59.2%
Terminal-Bench 3.026%15.7%34.6%34.1%
Harvey LAB (Vals)15.8%12.9%2.5%11.3%

The pattern against GPT-5.6 Sol Max: Grok 4.6 wins six of eight rows — including knowledge work (GDPVal-AA v2) and legal work (Harvey LAB, by a wide margin) — and loses the two hardest agentic-coding rows, DeepSWE v1.1 (65.9% vs 73%) and Terminal-Bench 3.0 (26% vs 34.6%). Against Fable 5 Max it wins GDPVal-AA v2 and Harvey LAB but trails on every coding row. SpaceXAI also cites 56.4% on APEX-SWE.

The independent read (Artificial Analysis)

Artificial Analysis titled its analysis “Grok 4.6 returns SpaceXAI to the intelligence frontier and leads on cost efficiency”. Its measurements:

MeasureGrok 4.6Context
Intelligence Index61Tied 3rd with GPT-5.6 Sol max; Opus 5 max 63, Fable 5 max 62, Kimi K3 just below
GDPVal-AA v21753 EloBehind only Claude Opus 5
τ³-Banking50.7%Top two, alongside Qwen3.8-Max (51.3%)
Terminal-Bench 2.188.4%In line with the leading models
AA-Briefcase1577 EloFable 5-tier; behind the Opus 5 family on long-horizon work
Cost per index task$0.84Lowest of any frontier model

Note the Terminal-Bench split: on the established 2.1 suite Artificial Analysis measures Grok 4.6 level with the leaders (88.4%), while on the newer, harder 3.0 version SpaceXAI’s own table shows it clearly behind (26% vs ~34% for both rivals). And the biggest gap in the record is what is absent: no SWE-bench Pro or Verified figure exists for Grok 4.6 in any harness — Grok 4.5 published 64.7% on SWE-Bench Pro; Grok 4.6 omits the benchmark family entirely, so there is no standardised repository-scale coding number to anchor on.

Pricing breakdown

Grok 4.6 keeps the $2/$6 sticker that made Grok 4.5 the value pick, with one significant catch above 200K input tokens.

ModeInput (per MTok)Output (per MTok)Cached input
Standard (≤200K input)$2.00$6.00$0.50
Long context (>200K input)$4.00$12.00$1.00
Grok 4.6 Fast$4.00$12.00

The long-context rate applies retroactively to every token in the request, not just those beyond 200K — a 210K-input call is billed at $4/$12 throughout. For workloads that routinely run large contexts, the effective price is closer to $4/$12 than the headline $2/$6.

Cost comparison with contemporaries

ModelInputOutputNotes
Grok 4.6$2.00$6.00Lowest cost per completed task on the frontier ($0.84, AA)
Kimi K3$3.00$15.00~$0.94 per AA task; open weights
Claude Opus 5$5.00$25.00Tops the AA index (63 max); ~2x the turns per task
GPT-5.6 Sol$5.00$30.00Ties Grok 4.6 on the AA index at 2.5x the output price
Fable 5$10.00$50.00The coding ceiling; 5x the input price
Qwen3.8-Max$2.00$6.00Price-matched; vendor-only benchmarks

How to access Grok 4.6

Via API

Grok 4.6 is generally available with no waitlist as grok-4.6 on SpaceXAI’s OpenAI-compatible API (console.x.ai), and through OpenRouter, Vercel and Cloudflare. EU availability was not confirmed at launch; Grok 4.5’s EU rollout also lagged the rest of the world.

from openai import OpenAI
client = OpenAI(base_url="https://api.x.ai/v1", api_key="XAI_API_KEY")

resp = client.chat.completions.create(
    model="grok-4.6",
    messages=[{"role": "user", "content": "Your prompt here"}],
)
print(resp.choices[0].message.content)

Via Cursor and the Grok apps

Grok 4.6 launched inside Cursor and Grok Build on day one, with 2x included usage for the first week. Grok Build access starts on the $30/month SuperGrok plan. Grok 4.6 also powers Grok Bot, SpaceXAI’s AI-teammate agent product launched in beta on 11 August 2026 for SuperGrok Heavy, Cursor Ultra and Cursor Teams Premium subscribers — see the Grok app page for detail.

TierPriceGrok 4.6Notes
Free$0Not confirmedUses lower Grok tiers
SuperGrok$30/moYesIncludes Grok Build
SuperGrok Heavy$300/moYesMulti-agent Heavy tier
Other tiersNot confirmedSuperGrok Lite / X Premium+ access unconfirmed at launch

See the Grok app page for the full consumer breakdown and the xAI provider page for company context.

How Grok 4.6 compares

Grok 4.6’s contemporaries are Claude Opus 5, Fable 5, GPT-5.6 Sol, Kimi K3 and its own five-week-old predecessor Grok 4.5.

vs GPT-5.6 Sol

The closest matchup on the board: the two models tie at 61 on the independent Artificial Analysis Intelligence Index. Grok 4.6 wins six of the eight published head-to-heads — most decisively knowledge work (GDPVal-AA v2: 1753 vs 1728) and legal work (Harvey LAB: 15.8% vs 2.5%) — at 2.5x less on input and 5x less on output ($2/$6 vs $5/$30). GPT-5.6 Sol Max keeps the two hardest agentic-coding benchmarks, DeepSWE v1.1 (73% vs 65.9%) and Terminal-Bench 3.0 (34.6% vs 26%), plus a published SWE-bench Pro figure (64.6% vendor) where Grok 4.6 has none — though Sol also carries METR’s reward-hacking flag, which Grok 4.6 so far does not. Choose Sol for maximum agentic-coding depth; choose Grok 4.6 for near-identical composite intelligence at a fraction of the cost.

vs Claude Opus 5

Claude Opus 5 tops the independent index (63 max vs 61) and leads the long-horizon AA-Briefcase eval, and it carries the strongest verified coding record short of Fable ($5/$25, 96.0% SWE-bench Verified, 79.2% SWE-bench Pro vendor). Grok 4.6’s counters are knowledge work — it is the only model within reach of Opus 5 on GDPVal-AA v2 — and efficiency: half the turns per task (~53 vs ~103) and roughly a quarter of the input-token volume. For verified repository-scale coding, Opus 5 is the safer pick; for high-volume agents and knowledge work per dollar, Grok 4.6.

vs Fable 5

Fable 5 Max stays one point ahead on the index (62 vs 61) and leads every coding row in SpaceXAI’s own table — CursorBench (70.5% vs 69.9%), FrontierCode (63.6% vs 61.3%), DeepSWE (70% vs 65.9%), APEX-Agents (59.2% vs 57.5%) — with the field’s best verified coding scores (80.3% SWE-bench Pro, 95.0% Verified, vendor). Grok 4.6 beats it on GDPVal-AA v2 and Harvey LAB, and costs a fifth as much ($2/$6 vs $10/$50). Fable 5 remains the ceiling; Grok 4.6 is the budget frontier alternative.

vs Grok 4.5

Grok 4.6 supersedes Grok 4.5 outright — the first Grok succession with no trade-off. Same price, same 500K window, same modalities, and better scores on all nine launch benchmarks: +5 on the AA index (61 vs 56), +227 Elo on GDPVal-AA v2, +11.9 points on DeepSWE v1.1, +10.4 on APEX-Agents. There is no remaining reason to prefer Grok 4.5; for video input or a 1M window, the older Grok 4.3 is still the only Grok that has either.

The practical consensus

Grok 4.6 is the first xAI model with an independently confirmed seat at the frontier table — tied with GPT-5.6 Sol on the composite, ahead of Kimi K3, behind only the Claude frontier pair — and it is the cheapest way to run frontier-level agents by a wide margin. What it is not, yet, is a verified coding leader: no SWE-bench number exists, and it loses the hardest agentic-coding evals to Sol and Fable. See our best AI models ranking, where Grok 4.6 debuts at #5, and best AI for coding for the coding-specific field.

Known limitations

No SWE-bench figure at all. Grok 4.5 published a SWE-Bench Pro score (64.7%); Grok 4.6 omits the benchmark family entirely, and no standardised harness has produced one. The most-cited repository-scale coding benchmark is simply absent from its record.

Weak on Terminal-Bench 3.0. SpaceXAI’s own table shows 26% against ~34% for GPT-5.6 Sol Max and Fable 5 Max — its clearest benchmark loss, partially offset by Artificial Analysis measuring 88.4% on the older Terminal-Bench 2.1.

Long-context calls double the price. Above 200K input tokens the rate becomes $4/$12 and applies to the whole request — workloads living near the 500K ceiling pay twice the headline price.

Same 500K window, no video. The context window and modalities are unchanged from Grok 4.5: half of Grok 4.3’s 1M window, text-and-image input only.

Vendor-selected benchmark slate. The composite index is independently confirmed, but the individual head-to-head rows use competitor numbers from system cards and leaderboards, on benchmarks SpaceXAI chose to publish.

EU availability unconfirmed. SpaceXAI did not confirm EU access at launch; Grok 4.5’s EU rollout lagged the rest of the world.

Company and safety baggage. SpaceXAI publishes less safety documentation than its peers, and the Grok apps’ image tools remain under scrutiny after the early-2026 deepfake scandal. See the xAI provider page.

Version history

VersionReleasedKey changes
Grok 4.612 Aug 2026Post-training release for long-running agents; independent AA index 61 (tied 3rd); self-verification behaviour; Fast variant; beats Grok 4.5 on all nine launch benchmarks
Grok 4.58 Jul 2026Coding/agentic focus, jointly trained with Cursor, 500K context, $2/$6 pricing; drops video input
Grok 4.3Apr 2026Native video input, document generation, 1M context, $1.25/$2.50
Grok 4.2Feb 2026Coding/reasoning push; among the first to clear 10% on ARC-AGI-2
Grok 4 / Grok 4 HeavyJul 2025Always-on reasoning; multi-agent Heavy variant

SpaceXAI is reported to be training Grok 5 on the Colossus 2 supercomputer.

Frequently asked questions

When was Grok 4.6 released?

Grok 4.6 launched on 12 August 2026, available the same day in Cursor, Grok Build and on SpaceXAI’s API, plus OpenRouter, Vercel and Cloudflare. It succeeded Grok 4.5 as xAI/SpaceXAI’s flagship model just five weeks after Grok 4.5 shipped.

How much does Grok 4.6 cost?

Via the API, Grok 4.6 is $2.00 per million input tokens and $6.00 per million output tokens, with cached input at $0.50/MTok. Requests with more than 200,000 input tokens are billed at a long-context rate of $4.00/$12.00 across the entire request, and the Grok 4.6 Fast variant costs double the standard rates. In the apps it is included from the $30/month SuperGrok plan.

Is Grok 4.6 as good as GPT-5.6 or Claude?

On the independent Artificial Analysis Intelligence Index, Grok 4.6 scores 61 — exactly tying GPT-5.6 Sol Max, one point behind Fable 5 Max (62) and two behind Claude Opus 5 max (63). It leads the field short of Opus 5 on knowledge work (GDPVal-AA v2: 1753) but trails both rivals on the hardest agentic-coding benchmarks and publishes no SWE-bench figure. It matches the frontier on composite intelligence; it does not lead it on coding.

What is Grok 4.6’s context window?

500,000 tokens, unchanged from Grok 4.5, with text and image input and text output. Note that requests above 200,000 input tokens move to the doubled long-context pricing tier.

How is Grok 4.6 different from Grok 4.5?

Grok 4.6 is a post-training upgrade on the same reported ~1.5T foundation, retuned with agentic reinforcement learning for long-running tasks — SpaceXAI reports emergent self-verification, with the model checking its own work between steps. It beats Grok 4.5 on all nine launch benchmarks (AA index 61 vs 56; GDPVal-AA v2 1753 vs 1526) at the same $2/$6 price, same 500K window and same modalities, so it supersedes Grok 4.5 with no trade-off.

Is Grok 4.6 the best-value frontier model?

On the evidence, yes. Artificial Analysis measures $0.84 per Intelligence Index task — the lowest of any frontier model, under Kimi K3 (~$0.94) and well under GPT-5.6 Sol and the Claude line — because Grok 4.6 completes agentic tasks in roughly half the turns of Claude Opus 5 (~53 vs ~103). The caveat: long-context requests (>200K input) double in price, and there is no verified SWE-bench coding figure. See best AI models for the full value picture.


Written at launch and last verified 13 August 2026. The Artificial Analysis Intelligence Index figure (61) is independent; individual head-to-head rows are from SpaceXAI’s launch table, which pulls competitor numbers from published system cards and leaderboards. Pricing and availability reflect the 12 August 2026 launch and are subject to change.