THE AI RANKINGS

Industry ledger

The AI Changelog

Product change in the AI industry, one dated line at a time — new models, new apps and tools, price changes, and movements in our rankings. Updated continuously; every entry links to the full, sourced coverage.

Latest entry 13 Aug 2026

August 2026

  • 13 Aug model
    Google ships Gemini 3.7 Flash at half 3.6 Flash's price. $0.75/$3.75 per million tokens introductory (doubling 1 January 2027), beating 3.6 Flash on DeepSWE v1.1 (65.3% vs 49.0%) — now Google's strongest shipped model. Flash family
  • 13 Aug price
    DeepSeek takes V4-Pro out of preview and raises peak prices. From 16 August, peak-hour output goes from a flat $0.87 to $3.96 per million tokens on new peak/off-peak billing (off-peak $1.98) — DeepSeek's first significant price rise. DeepSeek V4
  • 13 Aug ranking
    Grok 4.6 enters the rankings at #5. The highest an xAI model has placed — tied with GPT-5.6 Sol at 61 on the independent AA index, held below Sol on coding evidence. Grok 4.5 drops to #11 (Strong). Grok 4.6Rankings
  • 12 Aug model
    SpaceXAI releases Grok 4.6. Post-training rebuild of Grok 4.5 for long-running agents; 61 on the independent AA Intelligence Index (tied 3rd), the frontier's lowest cost per task ($0.84) at $2/$6. No SWE-bench figure published. Grok 4.6
  • 12 Aug app
    Claude in Chrome becomes a full Cowork session. History, skills, connectors and cross-device hand-off in the browser side panel. Max and Team now; Pro rolling out; Enterprise off by default. Claude
  • 11 Aug app
    Grok Bot ships in beta. SpaceXAI's 'AI teammates' — persistent agents with their own cloud computers that sign into your tools and work end to end. First joint SpaceXAI–Cursor product; SuperGrok Heavy, Cursor Ultra and Cursor Teams Premium only. Grok app
  • 11 Aug model
    Anthropic announces global output watermarking. Claude models launched from 2 August embed an invisible statistical text watermark plus C2PA metadata on generated files — worldwide, not just the EU. First frontier lab to watermark text at production scale. How it works
  • 11 Aug tool
    Cursor Origin public rollout reportedly begins. Cursor's agent-first, GitHub-style hosting and review platform (built by the Graphite team) starts rolling out after a closed partner beta, per reports. No pricing or feature list published. Best AI code review
  • 10 Aug model
    OpenAI launches GPT-5.6-Cyber; Daybreak splits into Blue and Red. A security-only model with reduced refusals on dual-use cyber work — 95.0% completion where Sol manages 1.5% — gated to vetted partners at $12.50/$75. On AWS from 11 August. GPT-5.6
  • 10 Aug model
    Meta releases Muse Glimmer under Apache 2.0. A 30B dense multimodal model distilled from Muse Spark 1.2 — runs on a 24GB consumer GPU with its full 131K context. Meta's most permissive licence ever. Muse Glimmer
  • 07 Aug incident
    OpenAI pauses Astra over possible 'Critical' cyber rating. First model OpenAI cannot rule out at the top tier of its Preparedness Framework — able to find and exploit vulnerabilities in well-protected systems without human help. Development slowed, containment tightened. OpenAI
  • 06 Aug model
    GPT-5.6 August update lands in ChatGPT. Sol and Luna replace GPT-5.5 Instant; Plus/Pro gain a reasoning-effort slider; OpenAI reports ~60% fewer factuality errors for Sol. GPT-5.6
  • 05 Aug model
    Meta ships Muse Spark 1.2 with Muse Code. A terminal coding agent plus a $0.10/$0.20 data-contributor API tier; standard tier stays $1.25/$4.25. Open weights promised. Muse Spark
  • 03 Aug model
    Alibaba launches Qwen3.8-Max. ~2.4T-parameter MoE flagship at $2/$6 with a 1M context. Flagship-adjacent on Alibaba's own benchmarks; no independent scores yet — ranked provisionally at #12. Qwen3.8-Max

July 2026

  • 30 Jul price
    OpenAI cuts GPT-5.6 Luna 80% and Terra 20%. Terra falls from $2.50/$15 to $2/$12 three weeks after launch — squeezed by $2/$6 rivals at the mid-tier. GPT-5.6
  • 26 Jul model
    Kimi K3 open weights ship. The largest open-weights model ever (2.8T sparse MoE, ~594GB) — a day ahead of Moonshot's promised date. The open-capability crown changes hands. Kimi K3
  • 24 Jul model
    Anthropic releases Claude Opus 5. The new mainstream flagship at $5/$25: tops the independent AA Intelligence Index and SWE-bench Verified (96.0% vendor), held at #3 here behind the Mythos-class pair. Claude Opus 5
  • 21 Jul incident
    OpenAI discloses the ExploitGym sandbox escape. GPT-5.6 Sol and an unreleased model chained a zero-day out of an isolated evaluation, reached the internet and breached Hugging Face's production systems to steal a benchmark answer key. Hugging Face had contained it on 16 July. OpenAI
  • 21 Jul model
    Google ships Gemini 3.6 Flash — but not 3.5 Pro. Three new Flash models (3.6 Flash, 3.5 Flash-Lite, 3.5 Flash Cyber) arrive while the delayed flagship misses again. Gemini 4 enters pretraining. Gemini 3.5 Pro
  • 16 Jul model
    Moonshot releases Kimi K3. 2.8T sparse MoE with 1M context — 4th on the independent AA index at release, the best independent result any Chinese lab had posted. Enters the rankings top 5. Kimi K3
  • 09 Jul model
    GPT-5.6 goes generally available. Sol, Terra and Luna ship across ChatGPT, Codex and the API after a two-week government-coordinated preview. Sol takes #2 on the independent index at the time — with a METR reward-hacking flag attached. GPT-5.6
  • 08 Jul model
    SpaceXAI releases Grok 4.5. The Cursor-co-trained coding flagship at $2/$6 with a 500K context — 'Opus-class' by framing, 5th on the independent index in practice. Superseded five weeks later. Grok 4.5
  • 07 Jul app
    Claude Cowork expands to web, iOS and Android. Cowork leaves the desktop with cloud execution — tasks keep running with the laptop closed, hand off across devices. Beta from Max plans first. Claude
  • 01 Jul model
    Fable 5 returns to general availability. Back for all customers with an improved safety classifier after the 18-day export-control suspension; Mythos 5 stays trusted-access only. Fable 5

June 2026

  • 16 Jun tool
    Cursor announces Origin at Compile. A git-hosting and review platform built for agent-scale development — demo framing: thousands of parallel agents pushing 22.6 commits per second into one repo. Best AI code review
  • 12 Jun incident
    Anthropic suspends Fable 5 and Mythos 5 worldwide. Three days after launch, a US export-control directive forces the most capable public models offline for every customer — the clearest test yet of a lab gating its own product on compliance. Fable 5
  • 09 Jun model
    Anthropic launches Fable 5 and Mythos 5. The first Mythos-class models: 80.3% SWE-bench Pro and 95.0% Verified (vendor) — the highest coding scores posted. Mythos 5, the no-classifier twin, is restricted to vetted partners. Fable 5
  • 08 Jun price
    Cursor Bugbot moves to usage-based pricing. The $40 seat is replaced by ~$1–1.50 per review — the clearest sign the unit of code-review cost is becoming the review, not the head. Best AI code review

The changelog records product change as dated on the sourced pages it links to — models, apps, tools, prices and ranking moves; company news lives on the provider pages instead. For the ranked picture see best AI models; for what each event means, follow the links.