Models
The models we track
The AI models behind ChatGPT, Claude, Gemini and others, most often in a task's top three first.
| Model | In the top three for | Cost / 1,000 requests | Evidence | Released |
|---|---|---|---|---|
| Claude Fable 5.1 Anthropic | 14 tasks business, maths and science, legal questions… | $55 | Independently tested | 1 Sept 2026 |
| Gemini 3.8 Flash Google DeepMind | 13 tasks studying and homework, planning and personal admin, business… | $4.13 | Independently tested | 2 Sept 2026 |
| GPT-6 Astra OpenAI | 7 tasks research, coding and building apps, Excel and spreadsheets… | $55 | Independently tested | 3 Sept 2026 |
| Claude Opus 5.5 Anthropic | 5 tasks planning and personal admin, translation and languages, writing… | $22 | Early results | 23 Sept 2026 |
| Kimi K3 Moonshot AI | 2 tasks studying and homework, maths and science | $17 | Independently tested | 17 Jul 2026 |
| GPT-5.6 Sol OpenAI | 2 tasks research, coding and building apps | $22 | Independently tested | – |
| Gemini 3.6 Flash Google DeepMind | 1 task research | $4.13 | Independently tested | 21 Jul 2026 |
| Gemini 3.1 Pro Google DeepMind | 1 task documents, PDFs and notes | $12 | Independently tested | – |
| GPT-6 Luna OpenAI | None yet | $0.55 | Early results | 22 Sept 2026 |
| GPT-6 Sol OpenAI | None yet | $11 | Early results | 22 Sept 2026 |
| DeepSeek V4.1 Flash DeepSeek | None yet | $0.72 | Independently tested | 10 Sept 2026 |
| DeepSeek V4 Pro DeepSeek | None yet | $1.48 | Independently tested | 13 Aug 2026 |
| Gemini 3.5 Flash-Lite Google DeepMind | None yet | $2.35 | Independently tested | 21 Jul 2026 |
| Claude Sonnet 5 Anthropic | None yet | $11 | Independently tested | 29 Jun 2026 |
| Claude Haiku 4.5 Anthropic | None yet | $5.50 | Independently tested | 15 Oct 2025 |
| GPT-5.6 Luna OpenAI | None yet | $1.24 | Independently tested | – |
Older versions
Replaced by a newer version. A new version starts with no results; we never assume it performs like the old one.