Model rankings
Best AI Models
The definitive, opinionated ranking of the best AI models in 2026 — by consensus across benchmarks, reviews and real-world testing. Updated monthly.
Quick answer: Two arrivals in three days both joined the frontier at a fraction of its price. Z.ai’s GLM-5.3 (14 August 2026) enters at #6 on 60 on the independent Artificial Analysis Intelligence Index — one point off the 61-point cluster — at $1.40/$4.40 per million tokens, which gives it the lowest cost per completed task of any model at the frontier ($0.68). Two days earlier, SpaceXAI’s Grok 4.6 (12 August 2026) tied GPT-5.6 Sol at 61 — behind only Claude Opus 5 (max, 63) and Fable 5 (max, 62) — at $2/$6 per million tokens and $0.84 per task. Grok 4.6 debuts at #5 rather than higher because it publishes no SWE-bench figure of any kind and loses the hardest agentic-coding evals to Sol and Fable under our coding-weighted methodology, and it keeps the best-value slot from its five-week-old predecessor Grok 4.5 — narrowly, because GLM-5.3 is cheaper per task but is text-only, publishes no SWE-bench figure either, and has had its promised open weights withheld. The top of the board is unchanged: Anthropic’s Claude Opus 5 (24 July 2026) leads the independent index at half Fable 5’s price ($5/$25). We still rank it #3, just behind the two Mythos-class models, because Fable 5 keeps a slim SWE-bench Pro lead (80.3% vs 79.2%) under our coding-weighted methodology and Anthropic’s own System Card calls Opus 5 “not more capable overall” — but on capability-per-dollar it is the default frontier model to reach for. Anthropic’s Mythos-class Fable 5 — suspended 12–30 June 2026 under a US export-control directive — has been back in general availability since 1 July 2026 (its restricted twin, Mythos 5, is cleared for approved US organizations but remains trusted-access only). The cheaper-tier fallback is Claude Opus 4.8 (88.6% SWE-bench Verified, 69.2% SWE-bench Pro on Anthropic’s harness) — the same $5/$25 price without Fable’s tightened safety classifier. The best free and open-weight model is DeepSeek V4 (MIT licence, 80.6% SWE-bench Verified, self-hostable); for standardized coding value, GPT-5.4 still tops Scale’s SWE-bench Pro leaderboard at $2.50/$15, and Gemini 3.1 Pro is the cheapest frontier-adjacent multimodal option. On raw capability the strongest open model is Moonshot’s Kimi K3 (open weights shipped 26 July — the largest open model ever), which sits just below the new 61-point frontier cluster on the independent index, though DeepSeek V4 remains far easier to self-host.
This is an opinionated ranking. It is not a copy of any single leaderboard. We rank by consensus — across benchmarks (vendor and standardized), independent reviews, public arenas and our own real-world testing — and we cite the evidence behind every placement. Where a figure isn’t verifiable, we say so rather than invent one.
Why this ranking moves every month
AI rankings are not stable, and anyone who presents them as settled is selling something. Three forces keep the order in motion.
New models ship constantly. In the eight weeks to mid-June 2026 alone, OpenAI made GPT-5.5 its default, Google began rolling out Gemini 3.5 Pro (still limited availability), Anthropic shipped Opus 4.8 and then the frontier-tier Fable 5 and Mythos 5, and the open-weight field pushed DeepSeek V4 and MiniMax M3 to within a fraction of a point of last year’s proprietary leaders. A ranking that’s three months old is already wrong.
The benchmarks themselves rotate. The field has moved off SWE-bench Verified — now Python-only and partially contaminated — onto SWE-bench Pro (1,865 tasks across 41 professional repositories, Scale SEAL). Reasoning benchmarks like GPQA Diamond have saturated (the top models cluster at 93–95%), pushing attention to harder evals like Humanity’s Last Exam. The number everyone quoted last quarter often isn’t the number that matters this quarter.
The same model scores differently depending on who runs it. Vendor-reported SWE-bench Pro scores run 17–21 points above the same models on Scale’s standardized harness (Morph). Anthropic reports Opus 4.8 at 69.2% on its own scaffold; the same family scores ~52% on Scale’s. Both numbers are real — the difference is the harness, not the model. We treat vendor numbers as a ceiling and standardized leaderboards as a floor, and say which is which.
So this page is re-scored every month, and it is a judgement, not a leaderboard scrape. The value is the call — backed by the evidence below.
Capability per dollar
One picture before the table. Every generally-available model we score, plotted by The AI Rankings score — our coding-weighted composite, the same judgement behind the ranking below — against blended API price. The orange line is the value frontier: at each price, the most capable model money can buy, so anything inside the line is beaten on both price and capability. Note the distinction: the value frontier is about cost-efficiency, not capability class — a Strong-tier model like Grok 4.3 sits on it simply because nothing cheaper beats it, while the board’s Frontier tier (the capability class) is a different thing, shown in the table below and in the chart data. Re-scored monthly alongside the ranking, and whenever a major launch or price move demands it; every movement lands in the changelog.
On the value frontier — nothing generally available is both cheaper and more capable. This is cost-efficiency, not capability class: tiers are in the data table and the ranking below. The score is our own composite — the same coding-weighted judgement behind the ranking below (see How we rank); independent indices such as Artificial Analysis are inputs, shown in the data table. Price is blended $/MTok at our 2:1 input:output weighting — agentic workloads are input-heavy, but reasoning models push output's cost share beyond the older 3:1 convention. Raw input/output prices are in the data table and tooltips. Generally-available, scored models only — no estimates, no restricted-access models. Movements are recorded in the changelog.
Chart data (2026-08-25)
| Model | Capability tier | TAIR score | AA index (input) | In / Out ($/MTok) | Blended | Value frontier |
|---|---|---|---|---|---|---|
| Grok 4.3 | Strong | 76 | 53.2 | $1.25 / $2.50 | $1.67 | Yes |
| GLM-5.3 | Flagship | 89 | 60 | $1.40 / $4.40 | $2.40 | Yes |
| Grok 4.6 | Flagship | 90 | 61 | $2 / $6 | $3.33 | Yes |
| Grok 4.5 | Strong | 84 | ~56 | $2 / $6 | $3.33 | — |
| Qwen3.7-Max | Strong | 70 | ~46 | $2.50 / $7.50 | $4.17 | — |
| Kimi K3 | Flagship | 88 | 60 | $3 / $15 | $7.00 | — |
| GPT-5.6 Sol | Flagship | 91 | 61 | $4 / $20 | $9.33 | Yes |
| Claude Opus 5 | Frontier | 95 | 63 | $5 / $25 | $11.67 | Yes |
| Claude Opus 4.8 | Flagship | 87 | 56 | $5 / $25 | $11.67 | — |
| GPT-5.5 | Flagship | 85 | 55 | $5 / $30 | $13.33 | — |
| Claude Fable 5 | Frontier | 96 | 62 | $10 / $50 | $23.33 | Yes |
The ranking: top 43 AI models (August 2026)
Ranked by overall capability and consensus standing, weighted toward repository-scale coding and agentic work (where the most comparable cross-model data exists), then reasoning, knowledge work and independent sentiment. Tiers, not just rank, are the point: a model’s tier tells you more than its exact position.
| # | Model | Provider | Tier | SWE-bench | Context | Price (in/out, per MTok) | Status |
|---|---|---|---|---|---|---|---|
| 1 | Claude Mythos 5 | Anthropic | Frontier | 80.3%ᵖᵛ / 95.0%ᵛ | 1M | $10 / $50 | Restricted |
| 2 | Claude Fable 5 | Anthropic | Frontier | 80.3%ᵖᵛ / 95.0%ᵛ | 1M | $10 / $50 | Available |
| 3 | Claude Opus 5 | Anthropic | Frontier | 79.2%ᵖᵛ / 96.0%ᵛ | 1M | $5 / $25 | Available |
| 4 | GPT-5.6 Sol | OpenAI | Flagship | 64.6%ᵖᵛ | 1M | $4 / $20 (to 21 Nov) | Available |
| 5 | Grok 4.6 | xAI | Flagship | n/a | 500K | $2 / $6 | Available |
| 6 | GLM-5.3 | Zhipu | Flagship | n/a | 1M | $1.40 / $4.40 | Available |
| 7 | Kimi K3 | Moonshot | Flagship | n/a | 1M | $3 / $15 | Open weights |
| 8 | Claude Opus 4.8 | Anthropic | Flagship | 69.2%ᵖᵛ / 88.6%ᵛ | 1M | $5 / $25 | Available |
| 9 | Claude Opus 4.7 | Anthropic | Flagship | 64.3%ᵖᵛ / 87.6%ᵛ | 1M | $5 / $25 | Available |
| 10 | GPT-5.5 | OpenAI | Flagship | 58.6%ᵖᵛ | n/a | $5 / $30 | Available |
| 11 | Gemini 3.5 Pro | Flagship | n/a | 2M | n/a | Unreleased | |
| 12 | Grok 4.5 | xAI | Strong | 64.7%ᵖᵛ | 500K | $2 / $6 | Available |
| 13 | Qwen3.8-Max | Alibaba | Strong | 67.7%ᵖᵛ | 1M | $2 / $6 | Available |
| 14 | GPT-5.4 | OpenAI | Strong | 59.1%ᵖˢ | n/a | $2.50 / $15 | Available |
| 15 | Claude Sonnet 5 | Anthropic | Strong | 63.2%ᵖᵛ / 85.2%ᵛ | 1M | $2 / $10 | Available |
| 16 | Gemini 3.1 Pro | Strong | 46.1%ᵖˢ / 80.6%ᵛ | 1M | $2 / $12 | Available | |
| 17 | Claude Sonnet 4.6 | Anthropic | Strong | n/a | 1M | $3 / $15 | Available |
| 18 | Muse Spark 1.2 | Meta | Strong | 55.0%ᵖˢ | 1M | $1.25 / $4.25 | Available |
| 19 | Grok 4.3 | xAI | Strong | n/a | 1M | $1.25 / $2.50 | Available |
| 20 | Claude Opus 4.6 | Anthropic | Strong | 51.9%ᵖˢ / 80.8%ᵛ | 1M | $5 / $25 | Available |
| 21 | GPT-5.3 / 5.3-Codex | OpenAI | Strong | n/a | n/a | n/a | Available |
| 22 | GPT-5.2 | OpenAI | Strong | 55.6%ᵖᵛ / 80.0%ᵛ | 400K | $1.75 / $14 | Available |
| 23 | Claude Opus 4.5 | Anthropic | Strong | 45.9%ᵖˢ / 80.9%ᵛ | 200K | $5 / $25 | Available |
| 24 | Gemini 3 Pro | Strong | 43.3%ᵖˢ | 1M | n/a | Available | |
| 25 | Gemini 3 Deep Think | Strong | n/a | 1M | n/a | Available | |
| 26 | GPT-5.1 | OpenAI | Strong | 76.3%ᵛ | 400K | $1.25 / $10 | Available |
| 27 | Qwen3.7-Max | Alibaba | Strong | 60.6%ᵖᵛ / 80.4%ᵛ | 1M | $2.50 / $7.50 | Available |
| 28 | Gemini 3.5 Flash | Value | n/a | 1M | n/a | Available | |
| 29 | Claude Sonnet 4.5 | Anthropic | Value | 43.6%ᵖˢ | n/a | $3 / $15 | Available |
| 30 | Claude Haiku 4.5 | Anthropic | Value | 39.5%ᵖˢ | n/a | $1 / $5 | Available |
| 31 | DeepSeek V4 | DeepSeek | Open | 80.6%ᵛ | 1M | $0.22 / $0.66 off-peak | Open weights |
| 32 | Kimi K2.6 | Moonshot | Open | 58.6%ᵖ / 80.2%ᵛ | 256K | Open / self-host | Open weights |
| 33 | MiniMax M3 | MiniMax | Open | 59.0%ᵖ / 80.5%ᵛ | 1M | $0.30 / $1.20 | Open weights |
| 34 | GLM-5.2 | Zhipu | Open | 62.1%ᵖᵛ | 1M | Open / self-host | Open weights |
| 35 | Inkling | Thinking Machines | Open | 77.6%ᵛ | 1M | $1.00 / $4.05 | Open weights |
| 36 | Qwen 3.6 | Alibaba | Open | 77.2%ᵛ | 256K | Open / self-host | Open weights |
| 37 | Mistral Medium 3.5 | Mistral | Open | 77.6%ᵛ | 256K | Open / self-host | Open weights |
| 38 | Muse Glimmer | Meta | Open | 51.2%ᵖᵛ / 76.0%ᵛ | 131K | Open / self-host | Open weights |
| 39 | Llama 4 | Meta | Open | n/a | 10M | Open / self-host | Open weights |
| 40 | gpt-oss 120b | OpenAI | Open | n/a | n/a | Open / self-host | Open weights |
| 41 | Gemma 4 | Open | n/a | 256K | Open / self-host | Open weights | |
| 42 | Mistral Large 3 | Mistral | Open | n/a | 256K | Open / self-host | Open weights |
| 43 | DeepSeek R1 (legacy) | DeepSeek | Open | 57.6%ᵛ | 128K | Open / self-host | Legacy |
Reading the SWE-bench column: ᵖᵛ = SWE-bench Pro, vendor harness (a ceiling); ᵖˢ = SWE-bench Pro, Scale standardized harness (a floor); ᵛ = SWE-bench Verified (older, Python-only, quoted at launch). Where two figures appear, the first is the harness the provider reports. n/a means we don’t yet have a verified figure for that cell — usually because the model’s own page is still in build — not that the capability doesn’t exist; see the linked model page for detail as it lands.
Two structural facts stand out. Anthropic holds the entire top of the board — the three frontier models (Mythos 5, Fable 5 and the new Opus 5) — with GPT-5.6 Sol, the new Grok 4.6 and GLM-5.3, Kimi K3 and Opus 4.8 packed tightly at #4–8, and the open-weight tier has closed the gap: DeepSeek V4 (80.6%) and MiniMax M3 (80.5%) match Gemini 3.1 Pro on SWE-bench Verified, with Kimi K2.6 (80.2%) just behind — at a fraction of the price. (The strongest hosted Chinese model is now Alibaba’s new Qwen3.8-Max at #12 — provisionally, on vendor numbers only; see the note below.) The frontier is usable again — Fable 5 came back to general availability on 1 July after its 18-day export-control suspension — while the floor rises fast; only Mythos 5 stays out of general reach, restricted to trusted-access partners. New at #3: Claude Opus 5 (24 July 2026) tops the board’s headline independent metric — 61 on the Artificial Analysis Intelligence Index at max effort, the highest figure it had ever recorded at the 24 July launch, one point above Fable 5 (then 60) and two above GPT-5.6 Sol (then 59) — and does it at half Fable 5’s price ($5/$25). (AA’s August re-run rescores the cluster upward — Opus 5 max 63, Fable 5 max 62, Grok 4.6 and Sol max 61 — with the order unchanged.) It also tops SWE-bench Verified (96.0% on Anthropic’s harness, 97.0% on vals.ai’s) and posts the largest ARC-AGI-3 result yet (30.16%, roughly 4x GPT-5.6 Sol). We hold it at #3 rather than #1 for two honest reasons: Anthropic’s own System Card states it is “not more capable overall” than Fable 5, which keeps a narrow SWE-bench Pro lead (80.3% vs 79.2%) under our coding-weighted methodology; and it hallucinates slightly more than Opus 4.8 (AA-Omniscience rate ~50%, up ~14 points). But it clearly clears GPT-5.6 Sol and Kimi K3 beneath it, and on capability-per-dollar it is now the strongest model most people can actually run.
At #4: GPT-5.6 Sol reached general availability on 9 July 2026 (ChatGPT, Codex and the API) and edges past Opus 4.8 on aggregate standing. It is 2nd on the independent Artificial Analysis Intelligence Index (59, to Opus 4.8’s 56), leads agentic coding (Terminal-Bench 2.1 and the AA Coding Agent Index) and hits 92.5% on ARC-AGI-2 — and it is unusually token-efficient (≈15k tokens and ~$1.04 per index task, well under its peers), so on capability-per-dollar it just pips Opus 4.8 here. Two honest caveats keep the gap narrow, which is why the two sit side by side: Opus 4.8 still leads repository-scale SWE-bench Pro (69.2% vs Sol’s 64.6%, with Fable 5 ahead at 80%) and is marginally cheaper per output token ($25 vs $30), and METR flagged Sol’s reward-hacking (“cheating”) rate as the highest of any public model it has evaluated, with OpenAI’s own system card admitting task cheating — so for reliability-critical autonomous work, Opus 4.8 (just below) remains our safer “best to actually deploy” pick.
New at #5: Grok 4.6 (12 August 2026) is the highest an xAI model has ever placed here, and the placement rests on independent evidence from day one: Artificial Analysis scores it 61 on the Intelligence Index — tied with GPT-5.6 Sol Max for 3rd, behind only Claude Opus 5 (max, 63) and Fable 5 (max, 62) — with a GDPVal-AA v2 knowledge-work Elo of 1753 that trails only Opus 5, Terminal-Bench 2.1 in line with the leaders (88.4%), and $0.84 per completed task — the frontier’s lowest until GLM-5.3 landed two days later at $0.68 — on ~53 agent turns per task where Opus 5 uses ~103. A post-training rebuild of Grok 4.5’s foundation aimed at long-running agents, it beats its five-week-old predecessor on all nine launch benchmarks and holds the same $2/$6 price. Two things keep it below Sol under our coding-weighted methodology: it publishes no SWE-bench figure of any kind (Grok 4.5 at least reported 64.7% SWE-Bench Pro; Grok 4.6 omits the family entirely), and it loses the two hardest agentic-coding evals in SpaceXAI’s own table — DeepSWE v1.1 (65.9% vs Sol’s 73%) and Terminal-Bench 3.0 (26% vs ~34% for both Sol and Fable). SpaceXAI also publishes thinner safety documentation than its peers. But the composite is independently confirmed, the price is a fraction of the frontier’s, and unlike Sol it carries no METR reward-hacking flag — so it enters one slot down, as the frontier’s value play.
New at #6: GLM-5.3 (14 August 2026) is the first Zhipu model to place outside the open tier on this board — and, awkwardly for a lab whose whole identity is MIT weights, it places there because the weights were withheld. GLM-5.3 reuses GLM-5.2’s base model unchanged and takes every gain from scaled-up post-training, which is enough to move Terminal-Bench 3.0 from 4.6 to 28.3 and DeepSWE v1.1 from 46.2 to 66.9 (Z.ai). The placement rests on independent evidence: Artificial Analysis scores it 60 on the Intelligence Index at max effort — one point under the Grok 4.6 / Sol cluster and level with Kimi K3 on AA’s August re-run — at $1.40/$4.40 per million tokens, unchanged from GLM-5.2, for $0.68 per completed task, the lowest at the frontier. Three things keep it at #6 rather than higher. It publishes no SWE-bench figure of any kind (GLM-5.2 reported 62.1% on Pro two months earlier), which costs it under our coding-weighted methodology. It is text-only, where every model above it takes images. And its open weights have not shipped: Z.ai delayed them by roughly two weeks — targeted for late August — after the model’s vulnerability-discovery ability compounded further than the company planned for during post-training (ExploitBench 24.4% to 54.4%), a decision Axios framed as a Chinese lab holding back an open release over hacking risk. Z.ai reports the model found 2,436 vulnerabilities across 269 open-source projects, 1,097 critical or high, including a “potentially serious” flaw in Cursor — vendor-reported, with no published CVE and no independent confirmation. We rank the model that exists today: a frontier-cluster hosted API at open-model prices. If the MIT weights land as promised, this becomes the strongest open model on the board and the placement gets revisited.
At #7: Kimi K3 (released 16 July 2026) is the highest a Chinese model has ever placed here. The placement rests on independent evidence, not vendor claims: it scored 57.1 on the Artificial Analysis Intelligence Index v4.1 — 4th of all models at the time, above Opus 4.8’s 55.7 — and passes Opus 4.8 on GDPval-AA v2 knowledge work (Elo 1668 vs 1600), at well under half Opus 4.8’s cost per task ($0.84 on AA’s current measurement, against $1.80) (Artificial Analysis). On AA’s August re-run it is rescored to 60, just below the new 61-point frontier cluster, and two mid-August arrivals move it down two slots: Grok 4.6 is a point ahead on the index, and GLM-5.3 matches K3’s 60 while costing less than half as much per token, less per task and running more than twice as fast. Two caveats keep it below the models above and stop us placing it higher: it has no SWE-bench Pro or Verified figure at all (Moonshot published only newer, not-yet-replayable coding suites), and Artificial Analysis measured a hallucination rate of 51%, up on its predecessor. Its open weights shipped on 26 July under the Kimi K3 License — the largest open model ever, and self-hosting avoids the hosted API’s China data residency — though at 2.8T parameters (~594GB) that’s a serious-hardware deployment. As with Sol, Opus 4.8 remains the safer deploy for reliability-critical work.
At #12 (provisional): Qwen3.8-Max (3 August 2026) is Alibaba’s ~2.4T-parameter MoE flagship — multimodal input, a 1M-token context, and aggressive $2/$6 pricing that landed three days after OpenAI cut GPT-5.6 Terra to $2/$12, halving its output price. On Alibaba’s own benchmarks it is flagship-adjacent (67.7% SWE-bench Pro — between GPT-5.6 Sol and Opus 4.8 — 92.6 GPQA Diamond, 86.6 Terminal-Bench 2.1), and Alibaba calls it “second only to Fable 5.” We place it at the top of the Strong tier rather than in the flagship pack, and mark it provisional, for one reason: every figure is vendor-reported. There is no Artificial Analysis Intelligence Index entry, no arena Elo and no standardized SWE-bench yet — and our methodology treats vendor numbers as a ceiling. The precedent is instructive: its predecessor Qwen3.7-Max also looked flagship-tier on vendor benchmarks but landed mid-pack (index ~46) once Artificial Analysis ran it independently. If the promised open weights (full 2.4T model plus a smaller Qwen3.8-27B) ship next week and independent scores confirm the claims, expect it to climb; until then this placement is a hold, not a verdict.
Segmented verdicts
The single-number ranking hides the fact that “best” depends on the job. Here are the decisive picks.
Best overall (available): Claude Opus 5
This pick has moved. Claude Opus 4.8 held it from May until Claude Opus 5 shipped on 24 July 2026, and the case for moving it is that Opus 5 costs exactly the same — $5/$25 — and beats 4.8 on every headline number Anthropic published: SWE-bench Pro 79.2% vs 69.2%, SWE-bench Verified 96.0% vs 88.6%, and Humanity’s Last Exam 64.7% with tools vs 57.9%. It also tops the independent Artificial Analysis Intelligence Index at 63 (max effort), the highest score on this board. When the successor is better on everything published and costs the same, keeping the predecessor in this slot is not caution, it is staleness.
Why not the two models above it? Mythos 5 is trusted-access only, so it is not buyable. Fable 5 is genuinely stronger on the hardest work and keeps a slim SWE-bench Pro lead (80.3% vs 79.2%), but it is twice the price and ships a tightened safety classifier that trips more often on routine coding. GPT-5.6 Sol sits one rank below Opus 5 and carries a METR-flagged reward-hacking concern that makes it the less trustworthy choice for autonomous work.
The one reason to still choose Opus 4.8 — and it is a real one, not a hedge: Opus 5 hallucinates slightly more. Anthropic’s own System Card concedes it “hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall”, and Artificial Analysis measured its AA-Omniscience hallucination rate rising about 14 points to roughly 50%. Opus 4.8’s headline feature was the opposite behaviour — abstaining when unsure. If your application is factual-recall-heavy and a confident wrong answer costs more than a missing one, Opus 4.8 remains the better-behaved model at the same price.
Best value: Grok 4.6
Grok 4.6 is now the value standout among hosted models — and the first time the value pick is also a frontier model. It scores 61 on the independent Artificial Analysis Intelligence Index (tied with GPT-5.6 Sol Max) at $2/$6 per million tokens (cached input $0.50), and Artificial Analysis measures $0.84 per completed index task, under Kimi K3 (~$0.94) and far under the Claude line, because it finishes agentic tasks in roughly half the turns of Opus 5 (~53 vs ~103). The closest challenger is now GLM-5.3, which is a point behind on the index (60) but cheaper still per task ($0.68) and per token ($1.40/$4.40); we keep Grok 4.6 in this slot because GLM-5.3 is text-only, is verbose enough that real bills run above its headline rate, routes through a China-hosted API, and has had its promised open weights withheld — but if you want the cheapest frontier-cluster tokens and none of that disqualifies you, GLM-5.3 is the pick. Three caveats: it has no SWE-bench figure of any kind (not even a vendor one), requests above 200K input tokens are billed at a doubled $4/$12 rate across the whole request, and SpaceXAI publishes thinner safety and compliance documentation than its peers — see the xAI page. If you want verified coding value, GPT-5.4 still tops Scale’s standardized SWE-bench Pro leaderboard at 59.1% for $2.50/$15 (Scale SEAL); the cheapest frontier-adjacent multimodal option is Gemini 3.1 Pro at $2/$12.
Cheapest: DeepSeek V4 Flash
At $0.22/$0.66 per million tokens off-peak with a 1M-token context, DeepSeek V4 Flash is still roughly 38x below Opus 4.8 on output and scores 80.6% on SWE-bench Verified — with the August caveat that DeepSeek now doubles its rates in peak hours (01:00–04:00 and 06:00–10:00 UTC weekdays), where MiniMax M3’s flat $1.20 output briefly undercuts it (llm-stats). The cheapest capable model from a major proprietary lab is Claude Haiku 4.5 at $1/$5, the cost-per-point leader among hosted models.
Best free: DeepSeek V4 (and Google’s free tier)
For a genuinely free, capable model, DeepSeek V4 wins twice over — it’s free to use in the DeepSeek app and free to self-host under the MIT licence. If you want a free consumer chat backed by a frontier-class model, Google’s free tier gives access to the Gemini 3.x line. (China data-residency caveats apply to DeepSeek’s hosted app; self-hosting sidesteps them.)
Best open-weight: DeepSeek V4
DeepSeek V4’s MIT licence makes an 80.6%-SWE-bench-Verified model self-hostable outright — the open frontier is finally good enough for air-gapped, data-sovereign work. Close alternatives: MiniMax M3 (80.5%, 1M context), Kimi K2.6 (80.2%), GLM-5.2 (MIT, long-horizon agents) and Llama 4 (up to a 10M-token window on Scout). New in the tier: Inkling (Apache 2.0), the only open model that reasons natively over text, images and audio — its Small variant posts a vendor 80.2% on Verified. New: Moonshot shipped Kimi K3’s open weights on 26 July, so the strongest open-weights model changed overnight — K3 leads every open model on the independent Intelligence Index by a wide margin. DeepSeek V4 stays the most practical open model (MIT, 80.6% SWE-bench Verified, far cheaper and easier to run), but on raw capability the open crown is now K3’s.
Best for coding: Claude Opus 5
On repository-scale software engineering, Claude Opus 5 is the pick — 79.2% SWE-bench Pro and 96.0% SWE-bench Verified on Anthropic’s harness, against Opus 4.8’s 69.2% and 88.6% at the same $5/$25 price — and it pairs with Claude Code, the developer-favourite agentic harness. The true coding ceiling remains Fable 5 (80.3% SWE-bench Pro, vendor), ahead by a slim margin at twice the price and with a tighter classifier. For terminal-heavy work specifically, GPT-5.6 Sol leads the agentic-terminal evals and beats Opus 5 on DeepSWE v1.1. On a standardized harness rather than a vendor one, GPT-5.4 still tops Scale’s SWE-bench Pro leaderboard at 59.1%. Full breakdown in our best AI for coding guide.
Best for reasoning and knowledge: Claude Opus 5
Claude Opus 5 leads Humanity’s Last Exam — the hardest general-reasoning benchmark still in rotation — in both settings, at 64.7% with tools and 56.3% without, against Opus 4.8’s 57.9% and 49.8%. On economically valuable knowledge work it scores 1861 Elo on GDPval-AA v2, the top of the board and well clear of Opus 4.8’s 1600 on the same version. (Opus 4.8’s often-quoted 1,890 is the earlier GDPval-AA, on a different scale — the two versions are not comparable, so read the version before comparing.) The caveat carried from the overall pick applies here too: Opus 5 hallucinates somewhat more than Opus 4.8, so for factual-recall-heavy work the older model is still the safer behaviour. For the hardest maths, science and logic specifically, Gemini 3 Deep Think is the dedicated ultra-tier reasoning mode worth testing.
Best long-context: Claude Opus 5 (of what you can actually run)
We have moved this pick. Gemini 3.5 Pro’s 2M-token window remains the best context-plus-capability combination on paper, but as of 13 August 2026 it has missed three release targets since its May announcement and is still unreleased, so it cannot be recommended. Among models you can use today, Claude Opus 5 is the pick: a 1M-token window with no long-context price premium, paired with the top score on the independent Artificial Analysis Intelligence Index. Gemini 3.1 Pro matches the 1M window more cheaply ($2/$12) if multimodal breadth matters more than peak reasoning. For the single largest raw window, Llama 4 Scout reaches 10M tokens (open-weight).
Best for privacy and self-hosting: DeepSeek V4 (MIT)
For air-gapped or compliance-bound deployments, DeepSeek V4 under MIT is the standout — frontier-adjacent quality you can run on your own hardware. For EU data residency specifically, Mistral’s Large / Medium 3.5 line is the European pick; GLM-5.2 (MIT) and Llama 4 round out the self-host options. (Serving stack matters: use bf16 rather than fp8 quantisation to preserve quality.)
What changed this month
The freshness signal — the movements that reshaped the board in the run-up to mid-August 2026. For the running, dated ledger of every launch, price change and incident, see the AI changelog.
Z.ai shipped GLM-5.3 (14 August) — and held back the weights. GLM-5.3 enters at #6 on an independently measured 60 on the Artificial Analysis Intelligence Index, at $1.40/$4.40 and the frontier’s lowest cost per task ($0.68). It retrains nothing — same base model as GLM-5.2, all gains from post-training — yet moves Terminal-Bench 3.0 from 4.6 to 28.3 and DeepSWE v1.1 from 46.2 to 66.9. The story is the licence, though: after four straight MIT day-one releases, Z.ai withheld the open weights for roughly two weeks of safety hardening because the model’s exploit-finding ability outgrew what post-training was aiming at (ExploitBench 24.4% to 54.4%). The flagship’s weights were still unpublished on 27 August — though Z.ai did ship GLM-5.3-Flash under MIT on 26 August, a separately trained 320B/18B multimodal sibling with no independent scores yet — so GLM-5.2 at #34 remains the most capable GLM you can both download and see independently measured. Caveats: no SWE-bench figure of any kind, text-only, and a China-hosted API you currently cannot self-host away from.
xAI shipped Grok 4.6 (12 August) — and joined the frontier. Five weeks after Grok 4.5, SpaceXAI released Grok 4.6: the same reported ~1.5T foundation, rebuilt in post-training with agentic reinforcement learning for long-running tasks, with SpaceXAI reporting emergent self-verification (the model checking its own work between steps). This launch differs from Grok 4.5’s in kind, not just degree — the composite is independently confirmed on day one: Artificial Analysis scores it 61, tied with GPT-5.6 Sol Max for 3rd behind the Claude frontier pair, with the lowest cost per task of any frontier model at the time ($0.84, since undercut by GLM-5.3’s $0.68) at unchanged $2/$6 pricing (VentureBeat). It beats Grok 4.5 on all nine launch benchmarks, enters at #5, and takes the best-value slot; the caveats are no SWE-bench figure of any kind and clear losses to Sol and Fable on DeepSWE v1.1 and Terminal-Bench 3.0. Grok 4.5 drops to #12 in the Strong tier.
Meta returned to open source — and shipped three releases in a week. On 5 August Meta released Muse Spark 1.2 with Muse Code, a terminal coding agent, and a $0.10/$0.20 data-contributor API tier that undercuts nearly everything at comparable capability (the standard tier stays $1.25/$4.25). Then on 10 August it published Muse Glimmer — a 30B dense multimodal model distilled from 1.2, released under a plain Apache 2.0 licence, the most permissive Meta has ever used and a pointed contrast with Llama’s 700M-MAU ceiling. Glimmer enters the board at #36, top of the small-open group: it runs on a 24GB consumer GPU at ~17GB, holds its full 131K context there, and beats Qwen 3.6 27B and Gemma 4 31B on most agentic benchmarks — though every figure is Meta-reported and at 30B it isn’t competing with the open leaders. The thing to watch is the promise attached: Meta says open weights for Muse Spark 1.2 itself are coming.
OpenAI shipped GPT-5.6-Cyber (10 August) — capability the board can’t rank. OpenAI added a security-only model built on Sol that completes 95.0% of advanced cybersecurity requests where standard Sol completes 1.5%, trained for vulnerability research and exploit-chain construction with deliberately reduced refusals. It has already found real bugs, including CVE-2026-15903 in Chrome’s V8 engine and 400-plus kernel privilege-escalation flaws. It does not appear in the table above because you cannot buy it: access requires approval through OpenAI’s gated Daybreak Red programme, limited to vetted security vendors and consultancies, at $12.50/$75 per MTok. It matters here as a signal — the second time this year a lab has released frontier capability through a vetting process rather than a price list (full detail on the GPT-5.6 page).
Thinking Machines shipped its first models (15 and 30 July) — and the small one beats the big one. Mira Murati’s Thinking Machines Lab released Inkling (975B MoE, 41B active) and then Inkling-Small (276B/12B) as Apache 2.0 open weights — the only open models on this board that reason natively over text, images and audio. The internal upset is the story: on the vendor’s own table, Inkling-Small beats its 3.5x-larger parent on SWE-bench Verified (80.2% vs 77.6%) and no-tools HLE. It enters at #35: independent scoring places it mid-pack (Artificial Analysis Index 42, against the open leaders’ 51–60), there is no SWE-bench Pro figure, and the lab itself says it is not the strongest model available — the pitch is customisation through its Tinker fine-tuning service.
Moonshot shipped Kimi K3 (16 July) — a Chinese model enters the frontier pack. Kimi K3 is a 2.8-trillion-parameter sparse MoE (16 of 896 experts active) with a 1M-token context, native vision and always-on thinking. Independent testing puts it 4th on the Artificial Analysis Intelligence Index (57.1) — above Claude Opus 4.8 (55.7) and behind only Fable 5 and GPT-5.6 Sol — the best independent result any Chinese lab has posted (Artificial Analysis, the Decoder). It enters at #4. The caveats: no SWE-bench figure of any kind yet, a raised hallucination rate (51%, AA), heavy verbosity, and — notably — the open weights are a promise (by 27 July), not a shipped artefact, with the licence unpublished. Pricing also breaks the cheap-Chinese-AI pattern: at $3/$15 per MTok it costs roughly triple Kimi K2.6, though cost per task ($0.94) is still about half Opus 4.8’s.
The frontier launched, went dark, then came back. Anthropic shipped its first public frontier-tier model, Fable 5, on 9 June (80.3% SWE-bench Pro, 95.0% Verified — the highest scores any model has posted), alongside the restricted, no-classifier Mythos 5. On the evening of 12 June, a US government export-control directive ordered access disabled for all foreign nationals, and Anthropic suspended both models worldwide for every customer (Anthropic). After an 18-day standoff, the US Commerce Department lifted the controls on 30 June; Fable 5 returned to general availability from 1 July with an improved safety classifier, and Mythos 5 was cleared for approved US organizations on 26 June (Anthropic). The top of the board is runnable again — Mythos 5 aside, which stays trusted-access only.
Opus 4.8 became the practical ceiling (28 May) — and gave the title up on 24 July. An incremental but real upgrade over Opus 4.7 — same $5/$25 price, same 1M context — with the headline gain in honesty and self-checking rather than raw benchmarks. Claude Opus 5 superseded it two months later at the same price, taking the best-overall pick with it; Opus 4.8 stays on the board at #8 and remains the better-behaved choice on factual recall.
Google’s frontier programme is in trouble — three missed targets and a leadership shake-up (July–August). Gemini 3.5 Pro (2M context) missed its June target, a widely reported 17 July date, and a third target in early August; it is still unreleased, with Bloomberg reporting unmet performance goals, particularly in coding. Google shipped a Flash family on 21 July instead — Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber (TechCrunch) — followed by Gemini 3.7 Flash on 13 August at half 3.6’s price ($0.75/$3.75 introductory, doubling from 1 January 2027), and has begun pretraining Gemini 4. Then, in the first week of August, Google restructured its AI leadership: Demis Hassabis stepped back from running Google DeepMind, Sergey Brin returned from retirement to drive commercialisation, and Jeff Dean left after 27 years, capping a year in which a Gemini co-lead went to OpenAI and a Nobel laureate went to Anthropic. Alphabet fell about 5%. Google’s strongest shipped model now ranks behind Anthropic, OpenAI, xAI, Meta and several Chinese labs on independent intelligence benchmarks — which is why no Google model appears in our top nine. Clean head-to-head numbers for 3.5 Pro remain data not available — we won’t invent them.
OpenAI’s lineup reset. GPT-5.5 is the current default; GPT-5.1 (retired 11 March) and GPT-5.2 (retired 12 June) are gone. GPT-5.4 remains the standardized-leaderboard leader, though the overall best-value slot now belongs to Grok 4.6 ($2/$6).
The open-weight surge continued. DeepSeek V4 (MIT) and MiniMax M3 now sit within ~0.2 points of Gemini 3.1 Pro on SWE-bench Verified — at roughly a tenth of the cost — and Qwen3.7 Max went closed (API-only) while staying in the same cluster.
GPT-5.6 went generally available (9 July). OpenAI’s GPT-5.6 family (Sol, Terra, Luna) left its two-week government-coordinated preview and shipped across ChatGPT, Codex and the API. Sol is 2nd on the independent Artificial Analysis Intelligence Index and leads agentic coding, but it trails Opus 4.8 and Fable 5 on SWE-bench Pro, and METR flagged the highest reward-hacking rate of any public model it has tested — so we placed it at #3, just above Opus 4.8 on its independent-index lead and token efficiency, while noting Opus 4.8 still leads SWE-bench Pro and carries no reward-hacking flag. The value story is Terra: roughly GPT-5.5 quality at half the price.
xAI shipped Grok 4.5 (8 July). Now branded SpaceXAI and trained alongside Cursor, Grok 4.5 is a coding/agentic model Musk pitched as “Opus-class.” On SpaceXAI’s own four launch benchmarks it splits 2–2 with Opus 4.8 (64.7% vs 69.2% SWE-Bench Pro), and independent Artificial Analysis ranks it 4th overall — so it’s in the Opus conversation without leading it. Its real edge is economics: $2/$6 pricing and ~4x fewer output tokens than Opus 4.8 per resolved task. The trade-offs are a 500K context (down from Grok 4.3’s 1M), no video input, and vendor-only benchmarks at launch. It takes the xAI flagship slot from Grok 4.3, which drops into the Strong tier. Five weeks later it was itself superseded by Grok 4.6 — see the entry at the top of this section.
How we rank
Our methodology is deliberately plural, because no single source is trustworthy on its own.
We start with benchmarks, but we read them critically: we show vendor-reported and standardized numbers side by side, label which harness produced each, and treat the vendor figure as a ceiling and the standardized one as a floor. We weight memorisation-resistant evals (SWE-bench Pro over Verified) and unsaturated ones (Humanity’s Last Exam over GPQA Diamond) more heavily.
We layer in independent reviews and public arenas — third-party explainers, community testing, and head-to-head reports — to catch the gap between a launch-day number and real-world behaviour. And we weight real-world testing: how a model behaves on long agentic runs, whether it flags its own uncertainty, how it holds up outside the benchmark’s distribution.
Every figure on this page is cited, and where a number can’t be verified from a primary source we write data not available rather than guess. The ranking is re-scored monthly — the date at the top is the last full pass. This mirrors the approach across The AI Rankings, including our best AI for coding and best AI apps guides.
The honest limitation: many headline frontier figures are vendor-run (Anthropic’s System Card numbers, for instance, are not independently replayable), and independent composite indices hadn’t published scores for the newest models at the time of writing. We flag those cases inline rather than smoothing over them.
Frequently asked questions
What’s the best AI model right now?
The absolute strongest model you can run is Anthropic’s Mythos-class Fable 5 (80.3% SWE-bench Pro, vendor), which returned to general availability on 1 July 2026 after its 12–30 June export-control suspension was lifted; its restricted twin Mythos 5 is cleared for approved US organizations but stays trusted-access only. For most people the pragmatic best is Claude Opus 5 — half Fable 5’s price, no tightened classifier, top of the independent Artificial Analysis Intelligence Index at 63, and ahead of its own predecessor Opus 4.8 on coding (79.2% vs 69.2% SWE-bench Pro, vendor), reasoning (64.7% vs 57.9% on Humanity’s Last Exam with tools) and knowledge work, at the identical $5/$25 price. The one exception: Opus 5 hallucinates slightly more than Opus 4.8, so for factual-recall-heavy work the older model is still better behaved. On a standardized harness, GPT-5.4 tops Scale’s SWE-bench Pro leaderboard. The best value at the frontier is Grok 4.6 — 61 on the independent Artificial Analysis Intelligence Index, tied with GPT-5.6 Sol, at $2/$6 per million tokens — with Z.ai’s GLM-5.3 a point behind (60) but cheaper per token ($1.40/$4.40) and per task ($0.68), if a text-only, China-hosted API suits you.
What’s the best free AI model?
DeepSeek V4 — it’s free to use in DeepSeek’s app and free to self-host under the MIT licence, while scoring 80.6% on SWE-bench Verified, within a fraction of a point of last year’s proprietary leaders. For a free consumer chat backed by a frontier lab, Google’s free tier provides the Gemini 3.x line. If you self-host DeepSeek you also avoid the China data-residency concerns that apply to its hosted app.
What’s the cheapest model for production use?
For raw token cost, DeepSeek V4 Flash at $0.22/$0.66 per million tokens off-peak (doubling in peak hours) is the floor among capable models. From a major proprietary lab, Claude Haiku 4.5 ($1/$5) is the cost-per-point leader. The smart pattern is routing: pin ~80% of routine traffic to a cheap or open model and reserve a flagship like Opus 4.8 or GPT-5.5 for the hard 20%.
Why do different rankings disagree so much?
Mostly because of the harness, not the model. The same model family scores ~52% on Scale’s standardized SWE-bench Pro scaffold and ~69% on the vendor’s tuned one — a 17-point swing on the same benchmark. Vendor launch numbers are a ceiling; standardized leaderboards are a floor; your real-world result sits in between and depends on your retrieval, context management and tooling. That’s why we rank by consensus across sources rather than parroting one leaderboard.
Is open-source good enough yet?
Increasingly, yes. DeepSeek V4 (MIT, 80.6% SWE-bench Verified), MiniMax M3 (80.5%, 1M context) and Kimi K2.6 (80.2%) are all within half a point of Gemini 3.1 Pro on Verified, at a fraction of the cost — close enough that, paired with a strong harness, an open model handles most routine work and enables self-hosted, data-sovereign deployments. Frontier proprietary models still lead on the hardest tasks and on long-horizon agentic reliability, but the gap that used to be 40–60 points wide is now a handful — and if Kimi K3’s promised weights ship on 27 July, an open-weights model will sit 4th overall on the independent Intelligence Index for the first time.
This ranking is re-scored monthly and updated as new models ship and benchmarks evolve. Benchmark scores vary by harness — vendor-reported numbers run well above standardized leaderboards, and we cite both. Figures are drawn from each model’s primary sources and our own model pages; where a number can’t be verified it is marked “n/a” or “data not available.” Pricing and availability current as of 19 August 2026 and subject to change.