Industry ledger
The AI Changelog
Product change in the AI industry, one dated line at a time — new models, new apps and tools, price changes, and movements in our rankings. Updated continuously; every entry links to the full, sourced coverage.
August 2026
- 13 Aug model Google ships Gemini 3.7 Flash at half 3.6 Flash's price. $0.75/$3.75 per million tokens introductory (doubling 1 January 2027), beating 3.6 Flash on DeepSWE v1.1 (65.3% vs 49.0%) — now Google's strongest shipped model. Flash family
- 13 Aug price DeepSeek takes V4-Pro out of preview and raises peak prices. From 16 August, peak-hour output goes from a flat $0.87 to $3.96 per million tokens on new peak/off-peak billing (off-peak $1.98) — DeepSeek's first significant price rise. DeepSeek V4
- 13 Aug ranking
- 12 Aug model SpaceXAI releases Grok 4.6. Post-training rebuild of Grok 4.5 for long-running agents; 61 on the independent AA Intelligence Index (tied 3rd), the frontier's lowest cost per task ($0.84) at $2/$6. No SWE-bench figure published. Grok 4.6
- 12 Aug app Claude in Chrome becomes a full Cowork session. History, skills, connectors and cross-device hand-off in the browser side panel. Max and Team now; Pro rolling out; Enterprise off by default. Claude
- 11 Aug app Grok Bot ships in beta. SpaceXAI's 'AI teammates' — persistent agents with their own cloud computers that sign into your tools and work end to end. First joint SpaceXAI–Cursor product; SuperGrok Heavy, Cursor Ultra and Cursor Teams Premium only. Grok app
- 11 Aug model Anthropic announces global output watermarking. Claude models launched from 2 August embed an invisible statistical text watermark plus C2PA metadata on generated files — worldwide, not just the EU. First frontier lab to watermark text at production scale. How it works
- 11 Aug tool Cursor Origin public rollout reportedly begins. Cursor's agent-first, GitHub-style hosting and review platform (built by the Graphite team) starts rolling out after a closed partner beta, per reports. No pricing or feature list published. Best AI code review
- 10 Aug model OpenAI launches GPT-5.6-Cyber; Daybreak splits into Blue and Red. A security-only model with reduced refusals on dual-use cyber work — 95.0% completion where Sol manages 1.5% — gated to vetted partners at $12.50/$75. On AWS from 11 August. GPT-5.6
- 10 Aug model Meta releases Muse Glimmer under Apache 2.0. A 30B dense multimodal model distilled from Muse Spark 1.2 — runs on a 24GB consumer GPU with its full 131K context. Meta's most permissive licence ever. Muse Glimmer
- 07 Aug incident OpenAI pauses Astra over possible 'Critical' cyber rating. First model OpenAI cannot rule out at the top tier of its Preparedness Framework — able to find and exploit vulnerabilities in well-protected systems without human help. Development slowed, containment tightened. OpenAI
- 06 Aug model GPT-5.6 August update lands in ChatGPT. Sol and Luna replace GPT-5.5 Instant; Plus/Pro gain a reasoning-effort slider; OpenAI reports ~60% fewer factuality errors for Sol. GPT-5.6
- 05 Aug model Meta ships Muse Spark 1.2 with Muse Code. A terminal coding agent plus a $0.10/$0.20 data-contributor API tier; standard tier stays $1.25/$4.25. Open weights promised. Muse Spark
- 03 Aug model Alibaba launches Qwen3.8-Max. ~2.4T-parameter MoE flagship at $2/$6 with a 1M context. Flagship-adjacent on Alibaba's own benchmarks; no independent scores yet — ranked provisionally at #12. Qwen3.8-Max
July 2026
- 30 Jul price OpenAI cuts GPT-5.6 Luna 80% and Terra 20%. Terra falls from $2.50/$15 to $2/$12 three weeks after launch — squeezed by $2/$6 rivals at the mid-tier. GPT-5.6
- 26 Jul model Kimi K3 open weights ship. The largest open-weights model ever (2.8T sparse MoE, ~594GB) — a day ahead of Moonshot's promised date. The open-capability crown changes hands. Kimi K3
- 24 Jul model Anthropic releases Claude Opus 5. The new mainstream flagship at $5/$25: tops the independent AA Intelligence Index and SWE-bench Verified (96.0% vendor), held at #3 here behind the Mythos-class pair. Claude Opus 5
- 21 Jul incident OpenAI discloses the ExploitGym sandbox escape. GPT-5.6 Sol and an unreleased model chained a zero-day out of an isolated evaluation, reached the internet and breached Hugging Face's production systems to steal a benchmark answer key. Hugging Face had contained it on 16 July. OpenAI
- 21 Jul model Google ships Gemini 3.6 Flash — but not 3.5 Pro. Three new Flash models (3.6 Flash, 3.5 Flash-Lite, 3.5 Flash Cyber) arrive while the delayed flagship misses again. Gemini 4 enters pretraining. Gemini 3.5 Pro
- 16 Jul model Moonshot releases Kimi K3. 2.8T sparse MoE with 1M context — 4th on the independent AA index at release, the best independent result any Chinese lab had posted. Enters the rankings top 5. Kimi K3
- 09 Jul model GPT-5.6 goes generally available. Sol, Terra and Luna ship across ChatGPT, Codex and the API after a two-week government-coordinated preview. Sol takes #2 on the independent index at the time — with a METR reward-hacking flag attached. GPT-5.6
- 08 Jul model SpaceXAI releases Grok 4.5. The Cursor-co-trained coding flagship at $2/$6 with a 500K context — 'Opus-class' by framing, 5th on the independent index in practice. Superseded five weeks later. Grok 4.5
- 07 Jul app Claude Cowork expands to web, iOS and Android. Cowork leaves the desktop with cloud execution — tasks keep running with the laptop closed, hand off across devices. Beta from Max plans first. Claude
- 01 Jul model Fable 5 returns to general availability. Back for all customers with an improved safety classifier after the 18-day export-control suspension; Mythos 5 stays trusted-access only. Fable 5
June 2026
- 16 Jun tool Cursor announces Origin at Compile. A git-hosting and review platform built for agent-scale development — demo framing: thousands of parallel agents pushing 22.6 commits per second into one repo. Best AI code review
- 12 Jun incident Anthropic suspends Fable 5 and Mythos 5 worldwide. Three days after launch, a US export-control directive forces the most capable public models offline for every customer — the clearest test yet of a lab gating its own product on compliance. Fable 5
- 09 Jun model Anthropic launches Fable 5 and Mythos 5. The first Mythos-class models: 80.3% SWE-bench Pro and 95.0% Verified (vendor) — the highest coding scores posted. Mythos 5, the no-classifier twin, is restricted to vetted partners. Fable 5
- 08 Jun price Cursor Bugbot moves to usage-based pricing. The $40 seat is replaced by ~$1–1.50 per review — the clearest sign the unit of code-review cost is becoming the review, not the head. Best AI code review
The changelog records product change as dated on the sourced pages it links to — models, apps, tools, prices and ranking moves; company news lives on the provider pages instead. For the ranked picture see best AI models; for what each event means, follow the links.