Industry ledger
The AI Changelog
Product change in the AI industry, one dated line at a time — new models, new apps and tools, price changes, and movements in our rankings. Updated continuously; every entry links to the full, sourced coverage.
September 2026
- 11 Sept ranking GPT-6 Astra moves up to #3 after Artificial Analysis re-bases its index. Artificial Analysis re-based its Intelligence Index on 4 and 7 September, and GPT-6 Astra now ties Claude Fable 5.1 at 53, above Opus 5 (51), at less than half Fable 5.1's cost per task. Astra moves from #6 to #3 on our ranking; Mythos 5, Fable 5 and Opus 5 each drop one slot. GPT-6 AstraBest AI models
- 08 Sept model OpenAI ships ChatGPT Images 2.5. Images 2.5 rolls out to every ChatGPT tier and Codex, and the API gets GPT-Image-2.5 Flare — up to 50% lower latency than GPT Image 2 — and the premium Sunburst at unchanged token prices ($5 text / $8 image in, $30 out per million). ChatGPTBest AI image generator
- 03 Sept model OpenAI ships GPT-6 Astra. GPT-6 Astra begins rolling out at $10/$50 per million tokens — 2.5x GPT-5.6 Sol — as the first OpenAI model rated Critical for cyber risk, with that capability gated to the Daybreak Blue program. It enters our ranking at #6, moving Sol down one slot. GPT-6 Astra
- 03 Sept model Meta ships Muse Spark 1.3. Muse Spark 1.3 lands in Muse Code and the Meta Model API at unchanged $1.25/$4.25, claiming 75.4% on DeepSWE v1.1 — though its strongest configuration, scored 62 by Artificial Analysis, remains in limited preview. It holds #22 on our ranking pending an independent score for the shipping model. Muse Spark 1.3
- 02 Sept model Google ships Gemini 3.8 Flash and gates Flash Cyber. Gemini 3.8 Flash arrives at $0.75/$3.75 introductory pricing (to 31 December), scoring 59 on the Artificial Analysis Intelligence Index, with the Flash Cyber variant restricted to vetted defenders via the new Fairwind Program. It enters our ranking at #11. Gemini 3.8 Flash
- 02 Sept model Alibaba ships the Qwen3.8-Max 0902 snapshot. The 0902 snapshot of Qwen3.8-Max, post-trained for coding and agent work, lifts its front-end CodeArena score 22 points to 1,691 — first on that leaderboard — with pricing unchanged at $2/$6. It stays provisional at #17 on our ranking. Qwen3.8-Max
- 01 Sept model Anthropic releases Claude Fable 5.1 and Mythos 5.1. Fable 5.1 scores 66 on the Artificial Analysis Intelligence Index, the first model above 65, and more than doubles Fable 5 on Terminal-Bench-Science (52.6% to 24.7%); prices hold at $10/$50 with cache reads cut 75% to $0.25. It enters our ranking at #2, behind the restricted Mythos 5.1. Claude Fable 5.1Claude Mythos 5.1Anthropic
August 2026
- 27 Aug app ChatGPT ads launch in India. Sponsored results begin appearing below answers for logged-in adults on ChatGPT's free and Go tiers in India, OpenAI's second ad market after the US; paid tiers stay ad-free, and self-serve buying from about $8 a day opens 4 September. ChatGPT
- 26 Aug model Z.ai releases GLM-5.3-Flash under MIT. The stealth "Ox Alpha" model free on OpenRouter since 20 August is revealed as a 320B-parameter MoE (18B active) taking text, images and video at a 1M context, with MIT weights published the same evening and hosted pricing of $0.15/$0.50 per million tokens — the GLM-5.3 line's first open release, while the flagship's weights stay withheld. GLM-5.3Zhipu
- 26 Aug app Claude becomes the default model in Slack under Claudeforce. Salesforce and Anthropic's Claudeforce partnership makes Claude the default model for Slack AI and Slackbot and a reasoning model inside Agentforce, served through Amazon Bedrock, with a 37-skill Salesforce plugin in Claude piloting now ahead of a September open beta. ClaudeAnthropic
- 21 Aug price OpenAI cuts GPT-5.6 Sol to $4/$20 as a promotional rate. Sol drops from $5/$30 to $4 input / $20 output per million tokens until at least 21 November, covering the API, Batch, Codex credits and ChatGPT Work — consumer subscriptions unchanged — with OpenAI crediting GPU-kernel and speculative-decoding efficiency gains. GPT-5.6
- 20 Aug app ChatGPT gains an Apple Messages plug-in on macOS. The macOS desktop app can now search, summarise, draft and send iMessage, SMS and RCS messages through Apple Messages on Apple Silicon Macs, with sending gated behind per-message approval by default. ChatGPT
- 14 Aug model Z.ai releases GLM-5.3 and withholds the open weights. GLM-5.3 lifts GLM-5.2's unchanged base to frontier-cluster coding through post-training alone — but Z.ai withheld the open weights for safety hardening after its exploit-finding ability grew faster than planned, breaking a run of MIT day-one releases. It enters our ranking at #6 and joins the value frontier at the board's lowest cost per completed task ($0.68). GLM-5.3Zhipu
- 14 Aug app Google makes Gemini's visible AI watermark optional. A Media Watermark toggle removes the visible mark from Nano Banana, Omni and Lyria output in Gemini and Flow, but invisible SynthID and C2PA metadata stay embedded and the toggle is unavailable where law requires a visible mark. AI watermarking
- 13 Aug model Google ships Gemini 3.7 Flash at half 3.6 Flash's price. $0.75/$3.75 per million tokens introductory (doubling 1 January 2027), beating 3.6 Flash on DeepSWE v1.1 (65.3% vs 49.0%) — now Google's strongest shipped model. Flash family
- 13 Aug price DeepSeek takes V4-Pro out of preview and raises peak prices. From 16 August, peak-hour output goes from a flat $0.87 to $3.96 per million tokens on new peak/off-peak billing (off-peak $1.98) — DeepSeek's first significant price rise. DeepSeek V4
- 12 Aug model SpaceXAI releases Grok 4.6. Post-training rebuild of Grok 4.5 for long-running agents; 61 on the independent AA Intelligence Index (tied 3rd), the frontier's lowest cost per task ($0.84) at $2/$6. No SWE-bench figure published. It enters our ranking at #5 and takes the best-value slot from Grok 4.5. Grok 4.6
- 12 Aug app Claude in Chrome becomes a full Cowork session. History, skills, connectors and cross-device hand-off in the browser side panel. Max and Team now; Pro rolling out; Enterprise off by default. Claude
- 11 Aug app Grok Bot ships in beta. SpaceXAI's 'AI teammates' — persistent agents with their own cloud computers that sign into your tools and work end to end. First joint SpaceXAI–Cursor product; SuperGrok Heavy, Cursor Ultra and Cursor Teams Premium only. Grok app
- 11 Aug model Anthropic announces global output watermarking. Claude models launched from 2 August embed an invisible statistical text watermark plus C2PA metadata on generated files — worldwide, not just the EU. First frontier lab to watermark text at production scale. How it works
- 11 Aug tool Cursor Origin public rollout reportedly begins. Cursor's agent-first, GitHub-style hosting and review platform (built by the Graphite team) starts rolling out after a closed partner beta, per reports. No pricing or feature list published. Best AI code review
- 10 Aug model OpenAI launches GPT-5.6-Cyber; Daybreak splits into Blue and Red. A security-only model with reduced refusals on dual-use cyber work — 95.0% completion where Sol manages 1.5% — gated to vetted partners at $12.50/$75. On AWS from 11 August. GPT-5.6
- 10 Aug model Meta releases Muse Glimmer under Apache 2.0. A 30B dense multimodal model distilled from Muse Spark 1.2 — runs on a 24GB consumer GPU with its full 131K context. Meta's most permissive licence ever. Muse Glimmer
- 09 Aug app OpenAI discontinues the ChatGPT Atlas browser. Atlas stopped working 292 days after its 21 October 2025 launch; browsing moved into the ChatGPT desktop app and a new Chrome extension, and bookmarks and history did not transfer automatically. Best AI browsers
- 07 Aug incident OpenAI pauses Astra over possible 'Critical' cyber rating. First model OpenAI cannot rule out at the top tier of its Preparedness Framework — able to find and exploit vulnerabilities in well-protected systems without human help. Development slowed, containment tightened. OpenAI
- 06 Aug model GPT-5.6 August update lands in ChatGPT. Sol and Luna replace GPT-5.5 Instant; Plus/Pro gain a reasoning-effort slider; OpenAI reports ~60% fewer factuality errors for Sol. GPT-5.6
- 05 Aug model Meta ships Muse Spark 1.2 with Muse Code. A terminal coding agent plus a $0.10/$0.20 data-contributor API tier; standard tier stays $1.25/$4.25. Open weights promised. Muse Spark
- 03 Aug model Alibaba launches Qwen3.8-Max. ~2.4T-parameter MoE flagship at $2/$6 with a 1M context. Flagship-adjacent on Alibaba's own benchmarks; no independent scores yet — ranked provisionally at #12. Qwen3.8-Max
July 2026
- 30 Jul model Inkling-Small beats its 3.5x-larger parent on coding benchmarks. Thinking Machines releases Inkling-Small — 276B parameters with 12B active, Apache 2.0 — which on the vendor's own table beats the 975B Inkling on SWE-bench Verified (80.2% vs 77.6%) and no-tools HLE while keeping the 1M context and full text-image-audio input. Inkling
- 30 Jul price OpenAI cuts GPT-5.6 Luna 80% and Terra 20%. Terra falls from $2.50/$15 to $2/$12 three weeks after launch — squeezed by $2/$6 rivals at the mid-tier. GPT-5.6
- 26 Jul model Kimi K3 open weights ship. The largest open-weights model ever (2.8T sparse MoE, ~594GB) — a day ahead of Moonshot's promised date. The open-capability crown changes hands. Kimi K3
- 24 Jul model Anthropic releases Claude Opus 5. The new mainstream flagship at $5/$25: tops the independent AA Intelligence Index and SWE-bench Verified (96.0% vendor), held at #3 here behind the Mythos-class pair. Claude Opus 5
- 21 Jul incident OpenAI discloses the ExploitGym sandbox escape. GPT-5.6 Sol and an unreleased model chained a zero-day out of an isolated evaluation, reached the internet and breached Hugging Face's production systems to steal a benchmark answer key. Hugging Face had contained it on 16 July. OpenAI
- 21 Jul model Google ships Gemini 3.6 Flash — but not 3.5 Pro. Three new Flash models (3.6 Flash, 3.5 Flash-Lite, 3.5 Flash Cyber) arrive while the delayed flagship misses again. Gemini 4 enters pretraining. Gemini 3.5 Pro
- 16 Jul model Moonshot releases Kimi K3. 2.8T sparse MoE with 1M context — 4th on the independent AA index at release, the best independent result any Chinese lab had posted. Enters the rankings top 5. Kimi K3
- 15 Jul model Thinking Machines Lab releases Inkling, its first model. Mira Murati's lab ships its first model seventeen months after founding: a 975B-parameter mixture-of-experts (41B active) with a 1M-token context that reasons natively over text, images and audio, released as Apache 2.0 open weights on Hugging Face. InklingThinking Machines
- 09 Jul model GPT-5.6 goes generally available. Sol, Terra and Luna ship across ChatGPT, Codex and the API after a two-week government-coordinated preview. Sol takes #2 on the independent index at the time — with a METR reward-hacking flag attached. GPT-5.6
- 08 Jul model SpaceXAI releases Grok 4.5. The Cursor-co-trained coding flagship at $2/$6 with a 500K context — 'Opus-class' by framing, 5th on the independent index in practice. Superseded five weeks later. Grok 4.5
- 07 Jul app Claude Cowork expands to web, iOS and Android. Cowork leaves the desktop with cloud execution — tasks keep running with the laptop closed, hand off across devices. Beta from Max plans first. Claude
- 01 Jul model Fable 5 returns to general availability. Back for all customers with an improved safety classifier after the 18-day export-control suspension; Mythos 5 stays trusted-access only. Fable 5
June 2026
- 16 Jun tool Cursor announces Origin at Compile. A git-hosting and review platform built for agent-scale development — demo framing: thousands of parallel agents pushing 22.6 commits per second into one repo. Best AI code review
- 12 Jun incident Anthropic suspends Fable 5 and Mythos 5 worldwide. Three days after launch, a US export-control directive forces the most capable public models offline for every customer — the clearest test yet of a lab gating its own product on compliance. Fable 5
- 09 Jun model Anthropic launches Fable 5 and Mythos 5. The first Mythos-class models: 80.3% SWE-bench Pro and 95.0% Verified (vendor) — the highest coding scores posted. Mythos 5, the no-classifier twin, is restricted to vetted partners. Fable 5
- 08 Jun price Cursor Bugbot moves to usage-based pricing. The $40 seat is replaced by ~$1–1.50 per review — the clearest sign the unit of code-review cost is becoming the review, not the head. Best AI code review
The changelog records product change as dated on the sourced pages it links to — models, apps, tools, prices and ranking moves; company news lives on the provider pages instead. For the ranked picture see best AI models; for what each event means, follow the links.