Grok 4.6
- Provider
- xAI
- Status
- Current
- Context
- 500,000 tok
- Price
- $2 / $6 /MTok
Grok 4.6 is xAI’s flagship model, launched on 12 August 2026 across Cursor, Grok Build and the SpaceXAI API, with partner availability on OpenRouter, Vercel and Cloudflare from day one. It succeeds Grok 4.5 after just five weeks, and it is built on the same reported ~1.5-trillion-parameter foundation — SpaceXAI attributes the gains almost entirely to post-training: curated model-generated reasoning data, regenerated supervised fine-tuning trajectories, and reinforcement learning on long agentic tasks spanning knowledge work, coding and domain-specific environments. The focus is long-running agents: SpaceXAI reports the model doing more self-testing and verification on extended trajectories, checking its own work before moving to the next step.
The launch claim is bigger than Grok 4.5’s — and this time the composite is independently confirmed. Artificial Analysis scores Grok 4.6 (high) at 61 on its Intelligence Index, tied with GPT-5.6 Sol Max for 3rd and one point behind Fable 5 Max (62), with Claude Opus 5 max on top (63) — the first time an xAI model has joined the frontier cluster, VentureBeat notes, overtaking Kimi K3. Its sharpest edge is economics: at $2/$6 per million tokens it posts the frontier’s lowest cost per task ($0.84 on Artificial Analysis’s index, using ~53 agent turns where Opus 5 uses ~103). The caveats are equally clear: it publishes no SWE-bench figure of any kind, and it loses to GPT-5.6 Sol Max on DeepSWE v1.1 and Terminal-Bench 3.0.
Quick specs
| Provider | xAI / SpaceXAI |
| Released | 12 August 2026 (Cursor, Grok Build, API, OpenRouter, Vercel, Cloudflare) |
| API model ID | grok-4.6 (OpenAI-compatible endpoint) |
| Context window | 500,000 tokens (unchanged from Grok 4.5) |
| Parameters | ~1.5 trillion (reported; same foundation as Grok 4.5) |
| Knowledge cutoff | Not disclosed |
| Input price | $2.00 / MTok ($4.00 above 200K input) |
| Output price | $6.00 / MTok ($12.00 above 200K input) |
| Cached input | $0.50 / MTok |
| Modalities | Text and image in; text out |
| AA Intelligence Index | 61 (independent; tied 3rd with GPT-5.6 Sol Max) |
| GDPVal-AA v2 | 1753 Elo (2nd, behind only Claude Opus 5) |
| Cost per AA index task | $0.84 — the lowest on the frontier |
| Best for | Cheap long-running agents, knowledge work, agentic coding at volume |
| Limitations | No SWE-bench figure at all; weak Terminal-Bench 3.0 (26%); long-context calls double the price; EU availability unconfirmed at launch |
What’s new in Grok 4.6
Grok 4.6 is a post-training release, not a new base model. SpaceXAI says it kept Grok 4.5’s foundation and changed the recipe: curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, an improved optimiser, regenerated SFT trajectories, and agentic reinforcement learning across knowledge work, general coding and domain-specific environments.
Built for long-running agents
Where Grok 4.5 was pitched at coding, Grok 4.6 is pitched at long agentic tasks — researching a topic, working across an entire codebase, or turning an idea into a working application over many iterations. The emergent behaviour SpaceXAI highlights is self-verification: on longer trajectories the model reviews its own work before moving on, a behaviour it attributes to reinforcement learning on long tasks rather than explicit instruction. The improvement over its predecessor is across the board — Grok 4.6 beats Grok 4.5 on all nine benchmarks in SpaceXAI’s launch table, including a 227-point jump on GDPVal-AA v2 (1753 vs 1526) and a 10-point jump on APEX-Agents (57.5% vs 47.1%).
The frontier’s cheapest completed task
Grok 4.6 keeps Grok 4.5’s headline economics and extends them. Artificial Analysis measures $0.84 per Intelligence Index task — under Kimi K3 ($0.94), GPT-5.6 Sol ($1.04) and far under the Claude line — driven by turn efficiency: ~53 agent turns and ~0.5B input tokens on average, versus ~103 turns and ~2.0B input tokens for Claude Opus 5. For long agentic runs, where input tokens compound every turn, that efficiency matters more than the sticker price.
Same window, same modalities
The trade-offs Grok 4.5 made are unchanged: a 500K context window (still half Grok 4.3’s 1M), text and image input only (no video), and text output. Reasoning effort is configurable, with “high” the setting behind the published benchmarks. New this release is a Grok 4.6 Fast variant at double the standard price for latency-sensitive work.
Benchmark performance
Two things distinguish this launch from Grok 4.5’s. First, the composite is independently verified on day one: Artificial Analysis’s own run puts Grok 4.6 at 61, exactly matching SpaceXAI’s table. Second, the vendor table’s competitor numbers are pulled from published system cards and leaderboards, not SpaceXAI reruns — more comparable than a tuned in-house harness, but still SpaceXAI’s selection of which benchmarks to show.
SpaceXAI’s launch table (12 August 2026)
| Benchmark | Grok 4.6 (high) | Grok 4.5 (high) | GPT-5.6 Sol Max | Fable 5 Max |
|---|---|---|---|---|
| AA Intelligence Index | 61 | 56 | 61 | 62 |
| GDPVal-AA v2 | 1753 | 1526 | 1728 | 1741 |
| CursorBench v3.2 | 69.9% | 66.7% | 67.2% | 70.5% |
| DeepSWE v1.1 | 65.9% | 54% | 73% | 70% |
| FrontierCode v1.1 | 61.3% | 56.6% | 60.6% | 63.6% |
| APEX-Agents | 57.5% | 47.1% | 56.7% | 59.2% |
| Terminal-Bench 3.0 | 26% | 15.7% | 34.6% | 34.1% |
| Harvey LAB (Vals) | 15.8% | 12.9% | 2.5% | 11.3% |
The pattern against GPT-5.6 Sol Max: Grok 4.6 wins six of eight rows — including knowledge work (GDPVal-AA v2) and legal work (Harvey LAB, by a wide margin) — and loses the two hardest agentic-coding rows, DeepSWE v1.1 (65.9% vs 73%) and Terminal-Bench 3.0 (26% vs 34.6%). Against Fable 5 Max it wins GDPVal-AA v2 and Harvey LAB but trails on every coding row. SpaceXAI also cites 56.4% on APEX-SWE.
The independent read (Artificial Analysis)
Artificial Analysis titled its analysis “Grok 4.6 returns SpaceXAI to the intelligence frontier and leads on cost efficiency”. Its measurements:
| Measure | Grok 4.6 | Context |
|---|---|---|
| Intelligence Index | 61 | Tied 3rd with GPT-5.6 Sol max; Opus 5 max 63, Fable 5 max 62, Kimi K3 just below |
| GDPVal-AA v2 | 1753 Elo | Behind only Claude Opus 5 |
| τ³-Banking | 50.7% | Top two, alongside Qwen3.8-Max (51.3%) |
| Terminal-Bench 2.1 | 88.4% | In line with the leading models |
| AA-Briefcase | 1577 Elo | Fable 5-tier; behind the Opus 5 family on long-horizon work |
| Cost per index task | $0.84 | Lowest of any frontier model |
Note the Terminal-Bench split: on the established 2.1 suite Artificial Analysis measures Grok 4.6 level with the leaders (88.4%), while on the newer, harder 3.0 version SpaceXAI’s own table shows it clearly behind (26% vs ~34% for both rivals). And the biggest gap in the record is what is absent: no SWE-bench Pro or Verified figure exists for Grok 4.6 in any harness — Grok 4.5 published 64.7% on SWE-Bench Pro; Grok 4.6 omits the benchmark family entirely, so there is no standardised repository-scale coding number to anchor on.
Pricing breakdown
Grok 4.6 keeps the $2/$6 sticker that made Grok 4.5 the value pick, with one significant catch above 200K input tokens.
| Mode | Input (per MTok) | Output (per MTok) | Cached input |
|---|---|---|---|
| Standard (≤200K input) | $2.00 | $6.00 | $0.50 |
| Long context (>200K input) | $4.00 | $12.00 | $1.00 |
| Grok 4.6 Fast | $4.00 | $12.00 | — |
The long-context rate applies retroactively to every token in the request, not just those beyond 200K — a 210K-input call is billed at $4/$12 throughout. For workloads that routinely run large contexts, the effective price is closer to $4/$12 than the headline $2/$6.
Cost comparison with contemporaries
| Model | Input | Output | Notes |
|---|---|---|---|
| Grok 4.6 | $2.00 | $6.00 | Lowest cost per completed task on the frontier ($0.84, AA) |
| Kimi K3 | $3.00 | $15.00 | ~$0.94 per AA task; open weights |
| Claude Opus 5 | $5.00 | $25.00 | Tops the AA index (63 max); ~2x the turns per task |
| GPT-5.6 Sol | $5.00 | $30.00 | Ties Grok 4.6 on the AA index at 2.5x the output price |
| Fable 5 | $10.00 | $50.00 | The coding ceiling; 5x the input price |
| Qwen3.8-Max | $2.00 | $6.00 | Price-matched; vendor-only benchmarks |
How to access Grok 4.6
Via API
Grok 4.6 is generally available with no waitlist as grok-4.6 on SpaceXAI’s OpenAI-compatible API (console.x.ai), and through OpenRouter, Vercel and Cloudflare. EU availability was not confirmed at launch; Grok 4.5’s EU rollout also lagged the rest of the world.
from openai import OpenAI
client = OpenAI(base_url="https://api.x.ai/v1", api_key="XAI_API_KEY")
resp = client.chat.completions.create(
model="grok-4.6",
messages=[{"role": "user", "content": "Your prompt here"}],
)
print(resp.choices[0].message.content)
Via Cursor and the Grok apps
Grok 4.6 launched inside Cursor and Grok Build on day one, with 2x included usage for the first week. Grok Build access starts on the $30/month SuperGrok plan. Grok 4.6 also powers Grok Bot, SpaceXAI’s AI-teammate agent product launched in beta on 11 August 2026 for SuperGrok Heavy, Cursor Ultra and Cursor Teams Premium subscribers — see the Grok app page for detail.
| Tier | Price | Grok 4.6 | Notes |
|---|---|---|---|
| Free | $0 | Not confirmed | Uses lower Grok tiers |
| SuperGrok | $30/mo | Yes | Includes Grok Build |
| SuperGrok Heavy | $300/mo | Yes | Multi-agent Heavy tier |
| Other tiers | — | Not confirmed | SuperGrok Lite / X Premium+ access unconfirmed at launch |
See the Grok app page for the full consumer breakdown and the xAI provider page for company context.
How Grok 4.6 compares
Grok 4.6’s contemporaries are Claude Opus 5, Fable 5, GPT-5.6 Sol, Kimi K3 and its own five-week-old predecessor Grok 4.5.
vs GPT-5.6 Sol
The closest matchup on the board: the two models tie at 61 on the independent Artificial Analysis Intelligence Index. Grok 4.6 wins six of the eight published head-to-heads — most decisively knowledge work (GDPVal-AA v2: 1753 vs 1728) and legal work (Harvey LAB: 15.8% vs 2.5%) — at 2.5x less on input and 5x less on output ($2/$6 vs $5/$30). GPT-5.6 Sol Max keeps the two hardest agentic-coding benchmarks, DeepSWE v1.1 (73% vs 65.9%) and Terminal-Bench 3.0 (34.6% vs 26%), plus a published SWE-bench Pro figure (64.6% vendor) where Grok 4.6 has none — though Sol also carries METR’s reward-hacking flag, which Grok 4.6 so far does not. Choose Sol for maximum agentic-coding depth; choose Grok 4.6 for near-identical composite intelligence at a fraction of the cost.
vs Claude Opus 5
Claude Opus 5 tops the independent index (63 max vs 61) and leads the long-horizon AA-Briefcase eval, and it carries the strongest verified coding record short of Fable ($5/$25, 96.0% SWE-bench Verified, 79.2% SWE-bench Pro vendor). Grok 4.6’s counters are knowledge work — it is the only model within reach of Opus 5 on GDPVal-AA v2 — and efficiency: half the turns per task (~53 vs ~103) and roughly a quarter of the input-token volume. For verified repository-scale coding, Opus 5 is the safer pick; for high-volume agents and knowledge work per dollar, Grok 4.6.
vs Fable 5
Fable 5 Max stays one point ahead on the index (62 vs 61) and leads every coding row in SpaceXAI’s own table — CursorBench (70.5% vs 69.9%), FrontierCode (63.6% vs 61.3%), DeepSWE (70% vs 65.9%), APEX-Agents (59.2% vs 57.5%) — with the field’s best verified coding scores (80.3% SWE-bench Pro, 95.0% Verified, vendor). Grok 4.6 beats it on GDPVal-AA v2 and Harvey LAB, and costs a fifth as much ($2/$6 vs $10/$50). Fable 5 remains the ceiling; Grok 4.6 is the budget frontier alternative.
vs Grok 4.5
Grok 4.6 supersedes Grok 4.5 outright — the first Grok succession with no trade-off. Same price, same 500K window, same modalities, and better scores on all nine launch benchmarks: +5 on the AA index (61 vs 56), +227 Elo on GDPVal-AA v2, +11.9 points on DeepSWE v1.1, +10.4 on APEX-Agents. There is no remaining reason to prefer Grok 4.5; for video input or a 1M window, the older Grok 4.3 is still the only Grok that has either.
The practical consensus
Grok 4.6 is the first xAI model with an independently confirmed seat at the frontier table — tied with GPT-5.6 Sol on the composite, ahead of Kimi K3, behind only the Claude frontier pair — and it is the cheapest way to run frontier-level agents by a wide margin. What it is not, yet, is a verified coding leader: no SWE-bench number exists, and it loses the hardest agentic-coding evals to Sol and Fable. See our best AI models ranking, where Grok 4.6 debuts at #5, and best AI for coding for the coding-specific field.
Known limitations
No SWE-bench figure at all. Grok 4.5 published a SWE-Bench Pro score (64.7%); Grok 4.6 omits the benchmark family entirely, and no standardised harness has produced one. The most-cited repository-scale coding benchmark is simply absent from its record.
Weak on Terminal-Bench 3.0. SpaceXAI’s own table shows 26% against ~34% for GPT-5.6 Sol Max and Fable 5 Max — its clearest benchmark loss, partially offset by Artificial Analysis measuring 88.4% on the older Terminal-Bench 2.1.
Long-context calls double the price. Above 200K input tokens the rate becomes $4/$12 and applies to the whole request — workloads living near the 500K ceiling pay twice the headline price.
Same 500K window, no video. The context window and modalities are unchanged from Grok 4.5: half of Grok 4.3’s 1M window, text-and-image input only.
Vendor-selected benchmark slate. The composite index is independently confirmed, but the individual head-to-head rows use competitor numbers from system cards and leaderboards, on benchmarks SpaceXAI chose to publish.
EU availability unconfirmed. SpaceXAI did not confirm EU access at launch; Grok 4.5’s EU rollout lagged the rest of the world.
Company and safety baggage. SpaceXAI publishes less safety documentation than its peers, and the Grok apps’ image tools remain under scrutiny after the early-2026 deepfake scandal. See the xAI provider page.
Version history
| Version | Released | Key changes |
|---|---|---|
| Grok 4.6 | 12 Aug 2026 | Post-training release for long-running agents; independent AA index 61 (tied 3rd); self-verification behaviour; Fast variant; beats Grok 4.5 on all nine launch benchmarks |
| Grok 4.5 | 8 Jul 2026 | Coding/agentic focus, jointly trained with Cursor, 500K context, $2/$6 pricing; drops video input |
| Grok 4.3 | Apr 2026 | Native video input, document generation, 1M context, $1.25/$2.50 |
| Grok 4.2 | Feb 2026 | Coding/reasoning push; among the first to clear 10% on ARC-AGI-2 |
| Grok 4 / Grok 4 Heavy | Jul 2025 | Always-on reasoning; multi-agent Heavy variant |
SpaceXAI is reported to be training Grok 5 on the Colossus 2 supercomputer.
Frequently asked questions
When was Grok 4.6 released?
Grok 4.6 launched on 12 August 2026, available the same day in Cursor, Grok Build and on SpaceXAI’s API, plus OpenRouter, Vercel and Cloudflare. It succeeded Grok 4.5 as xAI/SpaceXAI’s flagship model just five weeks after Grok 4.5 shipped.
How much does Grok 4.6 cost?
Via the API, Grok 4.6 is $2.00 per million input tokens and $6.00 per million output tokens, with cached input at $0.50/MTok. Requests with more than 200,000 input tokens are billed at a long-context rate of $4.00/$12.00 across the entire request, and the Grok 4.6 Fast variant costs double the standard rates. In the apps it is included from the $30/month SuperGrok plan.
Is Grok 4.6 as good as GPT-5.6 or Claude?
On the independent Artificial Analysis Intelligence Index, Grok 4.6 scores 61 — exactly tying GPT-5.6 Sol Max, one point behind Fable 5 Max (62) and two behind Claude Opus 5 max (63). It leads the field short of Opus 5 on knowledge work (GDPVal-AA v2: 1753) but trails both rivals on the hardest agentic-coding benchmarks and publishes no SWE-bench figure. It matches the frontier on composite intelligence; it does not lead it on coding.
What is Grok 4.6’s context window?
500,000 tokens, unchanged from Grok 4.5, with text and image input and text output. Note that requests above 200,000 input tokens move to the doubled long-context pricing tier.
How is Grok 4.6 different from Grok 4.5?
Grok 4.6 is a post-training upgrade on the same reported ~1.5T foundation, retuned with agentic reinforcement learning for long-running tasks — SpaceXAI reports emergent self-verification, with the model checking its own work between steps. It beats Grok 4.5 on all nine launch benchmarks (AA index 61 vs 56; GDPVal-AA v2 1753 vs 1526) at the same $2/$6 price, same 500K window and same modalities, so it supersedes Grok 4.5 with no trade-off.
Is Grok 4.6 the best-value frontier model?
On the evidence, yes. Artificial Analysis measures $0.84 per Intelligence Index task — the lowest of any frontier model, under Kimi K3 (~$0.94) and well under GPT-5.6 Sol and the Claude line — because Grok 4.6 completes agentic tasks in roughly half the turns of Claude Opus 5 (~53 vs ~103). The caveat: long-context requests (>200K input) double in price, and there is no verified SWE-bench coding figure. See best AI models for the full value picture.
Written at launch and last verified 13 August 2026. The Artificial Analysis Intelligence Index figure (61) is independent; individual head-to-head rows are from SpaceXAI’s launch table, which pulls competitor numbers from published system cards and leaderboards. Pricing and availability reflect the 12 August 2026 launch and are subject to change.