THE AI RANKINGS

Alibaba

Qwen3.8-Max

Provider
Alibaba
Status
Current
Context
1,000,000 tok
SWE-bench
67.7%
Price
$2 / $6 /MTok

Qwen3.8-Max is Alibaba’s new flagship Qwen model, released on 3 August 2026 after a preview on 19 July. It is a ~2.4-trillion-parameter mixture-of-experts reasoning model with a 1-million-token context window and, for the first time in the Max line, multimodal input (text, image and video in; text out). Alibaba positions it as its most capable model yet — claiming it is “second only to Fable 5” and beats GPT-5.6 Sol and Claude Fable 5 on several of its cited coding, general and multimodal benchmarks.

Two things make the launch notable beyond the benchmarks. First, price: at $2 input / $6 output per million tokens it landed three days after OpenAI cut its GPT-5.6 Terra tier to $2/$12, deliberately halving OpenAI’s output price at a similar capability tier. Second, open weights: Alibaba says the full 2.4T model — plus a smaller Qwen3.8-27B that runs on hardware you can actually buy — will be released as open weights the following week, which would make it the first Qwen Max-class model to go open and, if the claims hold, one of the most capable open models available. The important caveat for now: every benchmark is vendor-reported, and no independent leaderboard (Artificial Analysis, LMArena) or standardized SWE-bench has scored it yet.

Quick specs

ProviderAlibaba (Qwen)
Released3 August 2026 (preview 19 July)
StatusCurrent — closed-weight/API-only at launch, open weights announced for the following week
ArchitectureMixture-of-experts, ~2.4T total params (active count unconfirmed)
Context window1,000,000 tokens (~991K input, ~131K output, up to ~262K reasoning budget)
ModalitiesText, image and video in; text out
LicenceProprietary at launch (open weights promised; licence not yet specified)
Input price$2.00 / MTok
Output price$6.00 / MTok
SWE-bench Pro67.7% (vendor) — SWE-bench Verified not disclosed
Best forLong-context, agentic and multimodal work at aggressive pricing
LimitationsVendor-only benchmarks; China-hosted; open weights and 27B specs not yet shipped

VIEW QWEN3.8-MAX →

What Qwen3.8-Max is

Qwen3.8-Max is a reasoning-native, agentic flagship. It ships with an explicit thinking mode (reasoning budget up to ~262K tokens), function calling, structured outputs and tool use, and it targets long-horizon agentic work — Alibaba routes it through its Qoder agentic IDE and exposes OpenAI-, DashScope- and Anthropic-compatible endpoints so external harnesses like Claude Code, Cursor, Cline and Codex work by swapping the base URL and model ID.

Architecturally it is a mixture-of-experts model of about 2.4 trillion total parameters — the first Qwen Max-tier model confirmed as MoE and above a trillion parameters. Alibaba did not publish the activated-parameter count in the launch materials; some outlets report ~95B active per token, but that figure is unconfirmed, so treat it with caution. The 1-million-token context holds roughly 991K input tokens (about 983K with thinking on) and up to ~131K output. Unlike the text-only Qwen3.7-Max, this generation accepts image and video input alongside text, making it Qwen’s first multimodal model above a trillion parameters. The knowledge cutoff is not disclosed.

The headline strategic move is the reversal on openness. Where Qwen3.7-Max broke with tradition by shipping closed and API-only, Alibaba says Qwen3.8-Max will return to open weights — the full 2.4T model plus a smaller Qwen3.8-27B aimed at on-prem GPUs — the week after launch. As of 3 August those weights are not out yet, no licence has been confirmed (do not assume Apache 2.0 despite Qwen precedent), and the 27B sibling’s specs, context and hardware requirements are unpublished. The launch is widely read as Alibaba’s answer to Moonshot’s Kimi K3 and the broader Chinese open-weight race.

Benchmark performance

Alibaba published a broad vendor benchmark table. There is no independent verification yet — no Artificial Analysis Intelligence Index entry, no LMArena Elo, and no standardized (Scale) SWE-bench figure — so read the whole table as a ceiling.

BenchmarkQwen3.8-Max (vendor)Notes
SWE-bench Pro67.7%Alibaba harness; between GPT-5.6 Sol (64.6) and Opus 4.8 (69.2) on vendor figures
SWE-bench VerifiedNot disclosedOnly the Pro split was published
GPQA Diamond92.6Vendor-reported; the top field clusters at 92–95%
Terminal-Bench 2.186.6Between Opus 4.8/Fable 5 (84.6) and GPT-5.6 Sol max (88.8) on Alibaba’s chart
OSWorld-Verified86.1Agentic computer-use
Humanity’s Last Exam43.6Vendor-reported
Artificial Analysis IndexNot yet scoredNo independent entry as of 3 Aug 2026
LMArena (Text Elo)Not yet ratedNo blind-vote rating yet

The pattern is a frontier-adjacent agentic and coding model on vendor numbers, priced far below the US frontier. The honest caution is precedent: Qwen3.7-Max also looked flagship-tier on Alibaba’s own benchmarks (80.4% SWE-bench Verified) yet landed mid-pack (index ~46, ~#12) once Artificial Analysis ran it independently — and AA flagged it as verbose, which eroded its price advantage. Until an independent leaderboard scores Qwen3.8-Max, its true standing is unconfirmed. See best AI models for where it sits on our board and why.

Pricing and access

Qwen3.8-Max is priced at $2.00 input / $6.00 output per million tokens on Alibaba Cloud (Token Plan), with implicit cache reads at $0.25 — aggressive for a model of this class, and a deliberate shot at OpenAI, which had cut its GPT-5.6 Terra tier to $2/$12 three days earlier. Qwen matched the input price and halved the output price at a comparable capability tier. Preview access ran at up to ~90% off the standard rate.

The API model ID is qwen3.8-max, reachable through the international endpoint (dashscope-intl.aliyuncs.com) or the China endpoint, with OpenAI-, DashScope- and Anthropic-compatible modes. It is also wired into Alibaba’s Qoder and QoderWork agentic tools and the consumer Qwen app.

Open weights are promised but not yet shipped. Alibaba says the full 2.4T model and the smaller Qwen3.8-27B will be released the week after launch on Hugging Face and ModelScope; the licence has not been confirmed. Until then, hosted access routes through Alibaba Cloud (China jurisdiction), and the open Qwen 3.6 models remain the self-hostable alternative.

How Qwen3.8-Max compares

Known limitations

Vendor-only benchmarks — no independent leaderboard, arena or standardized SWE-bench has scored it yet, and the predecessor’s independent results came in well below its vendor claims. Open weights not yet shipped — the headline open-weight release (full model + Qwen3.8-27B) is promised for the following week but is not out, and no licence is confirmed. China-hosted — hosted access carries data-residency, compliance and content-control considerations. Undisclosed internals — activated-parameter count, knowledge cutoff, SWE-bench Verified and MMLU/AIME figures were not published. 27B sibling unspecified — beyond the name, its parameters, context, license, benchmarks and hardware requirements are unknown.

FAQ

What is Qwen3.8-Max?

Qwen3.8-Max is Alibaba’s flagship Qwen model, released 3 August 2026 — a ~2.4-trillion-parameter mixture-of-experts reasoning model with a 1-million-token context window and multimodal (text, image, video) input, built for long-horizon agentic and coding work.

Is Qwen3.8-Max open source?

Not yet, but Alibaba says it will be. It launched closed and API-only, with the full 2.4T open weights and a smaller Qwen3.8-27B announced for the following week on Hugging Face and ModelScope. The licence has not been confirmed, so don’t assume Apache 2.0 until Alibaba states it. For open Qwen you can run today, see the Qwen 3.6 models.

How much does Qwen3.8-Max cost?

$2.00 per million input tokens and $6.00 per million output on Alibaba Cloud, with cached input at $0.25. That deliberately undercut OpenAI, which had cut its GPT-5.6 Terra tier to $2/$12 three days earlier — Qwen matched the input price and halved the output price.

How good is Qwen3.8-Max?

On Alibaba’s own benchmarks, very strong — 67.7% SWE-bench Pro, 92.6 GPQA Diamond and 86.6 Terminal-Bench 2.1, which Alibaba says is “second only to Fable 5.” But every figure is vendor-reported, and no independent leaderboard has scored it yet, so treat those as a ceiling. Its predecessor looked flagship-tier on vendor numbers but ranked mid-pack once independently tested.

What is Qwen3.8-27B?

Qwen3.8-27B is a smaller open-weight sibling announced alongside the flagship, built to run on prosumer on-premise GPU hardware. Beyond the name, its exact specs, context window, licence and benchmarks have not been published as of 3 August 2026.

How does Qwen3.8-Max compare to Qwen3.7-Max?

Qwen3.8-Max is cheaper ($2/$6 vs $2.50/$7.50), adds image and video input (the Max line was previously text-only), posts stronger vendor coding and agentic scores, and is set to return to open weights — where Qwen3.7-Max shipped closed and API-only.


Last verified 3 August 2026. The 3 August launch, ~2.4T MoE architecture, 1M context, multimodal input, $2/$6 pricing and the announced open-weight release are drawn from Alibaba’s launch materials and multiple independent write-ups. All benchmark figures are vendor-reported (Alibaba’s own testing) — no Artificial Analysis, LMArena or standardized SWE-bench result exists yet — and the activated-parameter count, knowledge cutoff and Qwen3.8-27B specs are undisclosed. Confirm figures against independent leaderboards before relying on them.