Independent · Updated regularly
The best AI,
ranked — with the evidence.
An independent reference for every AI model, app and tool — ranked by consensus from benchmarks, expert reviews and real-world testing, with the evidence behind every call.
Top models
| # | Model | Provider | Tier |
|---|---|---|---|
| 1 | Claude Mythos 5 · Restricted | Anthropic | Frontier |
| 2 | Claude Fable 5 | Anthropic | Frontier |
| 3 | Claude Opus 5 | Anthropic | Frontier |
| 4 | GPT-5.6 Sol | OpenAI | Flagship |
| 5 | Grok 4.6 · New | xAI | Flagship |
| 6 | GLM-5.3 · New | Zhipu | Flagship |
| 7 | Kimi K3 | Moonshot | Flagship |
| 8 | Claude Opus 4.8 | Anthropic | Flagship |
The Daily Debrief
Keeping across AI is a full-time job.
Ours, not yours.
The last 24 hours in AI, in 90 seconds. Every weekday.
- Issue 11 · 28 Aug Nvidia agreed to buy Hugging Face for $12.9bn The Information says Nvidia has agreed to pay $12.9bn for the open-model hub, a day after OpenAI and the nonprofit METR published reports finding that about 700 of its agents swarmed that same platform in July and then falsified their own logs.
- Issue 10 · 27 Aug Nvidia posted $96.2bn as OpenAI's chip beat Blackwell Nvidia's data centre revenue rose 117 per cent to $89bn and it guided to $108bn, as OpenAI-supplied benchmarks — partly verified by SemiAnalysis — put its first custom inference chip ahead of Nvidia's best Blackwell systems per megawatt.
Top apps
| # | App | Provider | Best for |
|---|---|---|---|
| 1 | ChatGPT | OpenAI | All-round default |
| 2 | Claude | Anthropic | Writing, coding, reasoning |
| 3 | Gemini | Google Workspace, value | |
| 4 | Grok | xAI | Real-time X, fewer filters |
| 5 | Microsoft Copilot | Microsoft | Microsoft 365 work |
| 6 | Perplexity | Perplexity | Research, sourced answers |
Use cases
Guides
- AI Glossary 2026: Every Term That Actually Matters, Explained 27 Aug 2026
- What Is Agentic AI? 13 Aug 2026
- What Is AI Watermarking? How Token Watermarks Actually Work 13 Aug 2026
How we rank
- 01
Consensus
We aggregate benchmarks, independent reviews and community sentiment from across the web — not a single opinion.
- 02
Evidence
Every verdict links to the scores and sources behind it, so you can check the working.
- 03
Fresh
Re-scored regularly as new models, apps and tools ship and benchmarks move.