PodcastIntelligence Snacks · Weekly conversations about AI, software and business, from the team behind Normie Mode.Listen →

Which AI is best right now?

Pick your task to see which AI does it best, where to get it, and the evidence behind our pick.

Updated 26 Sept 2026

Best AI for…

The everyday jobs people search for, most popular first, with the models that do each best.

All 46 tasks →

Already paying for one?

How the best model on each plan does across the 15 everyday tasks above.

App and planPriceIn the top three for
ChatGPT FreeFree0 of 15 tasks
ChatGPT Go$8/month0 of 15 tasks
ChatGPT Plus$20/month7 of 15 tasks
coding and building apps, business, research, accounting and finance…
ChatGPT Pro$100/month7 of 15 tasks
coding and building apps, business, research, accounting and finance…
Claude FreeFree0 of 15 tasks
Claude Pro$20/month5 of 15 tasks
writing, cvs, job applications and interviews, health questions, planning and personal admin…
Claude Max$100/month14 of 15 tasks
coding and building apps, writing, studying and homework, cvs, job applications and interviews…
Gemini FreeFree2 of 15 tasks
research, documents, pdfs and notes
Google AI Plus$4.99/month2 of 15 tasks
research, documents, pdfs and notes
Google AI Pro$19.99/month13 of 15 tasks
writing, studying and homework, cvs, job applications and interviews, business…
Google AI Ultra$99.99/month13 of 15 tasks
writing, studying and homework, cvs, job applications and interviews, business…
Kimi FreeFree2 of 15 tasks
studying and homework, maths and science
Kimi Moderato$19/month2 of 15 tasks
studying and homework, maths and science

Compare the models

The models behind the apps, side by side: how capable, how much people like them and what they cost developers.

Most capable

GPT-6 Astra and Claude Fable 5.1 are the most capable overall, followed by GPT-5.6 Sol and Kimi K3.

People's favourites

Claude Opus 5.5 and Claude Fable 5.1 are preferred most in blind votes, followed by Gemini 3.8 Flash and Gemini 3.1 Pro.

Cheapest

GPT-6 Luna and DeepSeek V4.1 Flash cost the least to use, followed by GPT-5.6 Luna and DeepSeek V4 Pro.

Reads the most

GPT-6 Astra and GPT-5.6 Sol can take in the most text at once, followed by GPT-5.6 Luna and GPT-6 Luna.

Too new to judge

Claude Opus 5.5, GPT-6 Luna and GPT-6 Sol came out in the last few days and have only 1–2 independent results so far.

Capability

One score that combines dozens of independent tests. The thin line on each bar is the likely range.

Overall capability

Position on Epoch AI's score, which combines dozens of independent tests. The thin line shows the likely range · Higher is better

If two models' ranges overlap, treat them as roughly equal.

GPT-6 Astra · OpenAI1st
Claude Fable 5.1 · Anthropic2nd
GPT-5.6 Sol · OpenAI3rd
Kimi K3 · Moonshot AI4th
Gemini 3.8 Flash · Google DeepMind5th
Claude Sonnet 5 · Anthropic6th
GPT-5.6 Luna · OpenAI7th
DeepSeek V4 Pro · DeepSeek8th
DeepSeek V4.1 Flash · DeepSeek9th
Gemini 3.1 Pro · Google DeepMind10th
Gemini 3.6 Flash · Google DeepMind11th
Gemini 3.5 Flash-Lite · Google DeepMind12th
Claude Haiku 4.5 · Anthropic13th
Claude Opus 5.5 · AnthropicNot tested yet
GPT-6 Luna · OpenAINot tested yet
GPT-6 Sol · OpenAINot tested yet
Source: Epoch AI · CC BY 4.0

Capability vs cost

Up and to the left is better: more capable for less money.

Capability against cost

Epoch capability score against list price per 1,000 typical requests (log scale)

How capable each model is against what it costsScatter chart. Horizontal: cost per 1,000 typical requests, log scale. Vertical: Epoch capability score with uncertainty range. The same data is in the table below.140145150155160165170$1.0$3.0$10$30Cost per 1,000 typical requests (list price) → cheaper on the leftCapability score (Epoch AI) → betterGPT-6 AstraClaude Fable 5.1GPT-5.6 SolKimi K3Gemini 3.8 FlashClaude Sonnet 5GPT-5.6 LunaDeepSeek V4 ProGemini 3.1 ProGemini 3.6 FlashGemini 3.5 Flash-LiteClaude Haiku 4.5
The vertical line on each dot is the likely range: if two lines overlap, it's too close to call. Sources: Epoch AI, models.dev

People's votes

People compare two anonymous answers and vote for the better one.

LMArena text ranking

Position among the models here, from votes as of 25 Sept 2026 · Higher is better

New models appear here after they collect enough votes, usually within a couple of weeks.

Claude Opus 5.5 · Anthropic1st
Claude Fable 5.1 · Anthropic2nd
Gemini 3.8 Flash · Google DeepMind3rd
Gemini 3.1 Pro · Google DeepMind4th
Gemini 3.6 Flash · Google DeepMind5th
Kimi K3 · Moonshot AI6th
DeepSeek V4.1 Flash · DeepSeek7th
GPT-5.6 Sol · OpenAI8th
DeepSeek V4 Pro · DeepSeek9th
GPT-6 Astra · OpenAI10th
Claude Sonnet 5 · Anthropic11th
Gemini 3.5 Flash-Lite · Google DeepMind12th
GPT-5.6 Luna · OpenAI13th
Claude Haiku 4.5 · Anthropic14th
GPT-6 Luna · OpenAINot tested yet
GPT-6 Sol · OpenAINot tested yet
Source: LMArena · CC BY 4.0

Cost

What you'd pay the maker directly. Chat apps with monthly plans price differently.

Cost per 1,000 typical requests

List price, assuming about 1,500 words in, 500 out per request. Models that think longer cost more in practice · Lower is better

GPT-6 Luna · OpenAI$0.55
DeepSeek V4.1 Flash · DeepSeek$0.72
GPT-5.6 Luna · OpenAI$1.24
DeepSeek V4 Pro · DeepSeek$1.48
Gemini 3.5 Flash-Lite · Google DeepMind$2.35
Gemini 3.8 Flash · Google DeepMind$4.13
Gemini 3.6 Flash · Google DeepMind$4.13
Claude Haiku 4.5 · Anthropic$5.50
Claude Sonnet 5 · Anthropic$11
GPT-6 Sol · OpenAI$11
Gemini 3.1 Pro · Google DeepMind$12
Kimi K3 · Moonshot AI$17
GPT-5.6 Sol · OpenAI$22
Claude Opus 5.5 · Anthropic$22
GPT-6 Astra · OpenAI$55
Claude Fable 5.1 · Anthropic$55
Source: models.dev · MIT

How much it can read

How much text it can take in at once, in pages of about 500 words.

Pages it can read at once

Based on each model's context window · Higher is better

GPT-6 Astra · OpenAI1,575
GPT-5.6 Sol · OpenAI1,575
GPT-5.6 Luna · OpenAI1,575
GPT-6 Luna · OpenAI1,575
GPT-6 Sol · OpenAI1,575
Kimi K3 · Moonshot AI1,573
Gemini 3.8 Flash · Google DeepMind1,573
Gemini 3.1 Pro · Google DeepMind1,573
Gemini 3.6 Flash · Google DeepMind1,573
Gemini 3.5 Flash-Lite · Google DeepMind1,573
Claude Fable 5.1 · Anthropic1,500
Claude Sonnet 5 · Anthropic1,500
DeepSeek V4 Pro · DeepSeek1,500
DeepSeek V4.1 Flash · DeepSeek1,500
Claude Opus 5.5 · Anthropic1,500
Claude Haiku 4.5 · Anthropic300
Source: models.dev · MIT

Too new to judge

What the makers say, until independent results arrive.

Anthropic · 23 Sept 2026

Claude Opus 5.5

The maker says“It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.”

Independent results so far: 2

OpenAI · 22 Sept 2026

GPT-6 Luna

The maker says“At high effort, GPT-6 Luna improves on its predecessor by 5.4 percentage points at 58% lower cost per task.”

Independent results so far: 1

OpenAI · 22 Sept 2026

GPT-6 Sol

The maker says“On AutomationBench, GPT-6 Sol at xhigh effort outperforms Claude Opus 5 at max effort at just 9% of Opus 5's cost per task.”

Independent results so far: 2

All the numbers

Everything above in one table.

ModelCapabilityPeople's votesCost / 1,000 requestsPages at onceEvidence
GPT-6 Astra
OpenAI
167
163–172
10th of 14
7,185 votes
$55
$10 / $50 per M tokens
1,575Independently tested
27 results
Claude Fable 5.1
Anthropic
165
162–170
2nd of 14
9,942 votes
$55
$10 / $50 per M tokens
1,500Independently tested
25 results
GPT-5.6 Sol
OpenAI
162
160–166
8th of 14
34,258 votes
$22
$4 / $20 per M tokens
1,575Independently tested
29 results
Kimi K3
Moonshot AI · open weights
158
155–161
6th of 14
26,400 votes
$17
$3 / $15 per M tokens
1,573Independently tested
24 results
Gemini 3.8 Flash
Google DeepMind
157
155–161
3rd of 14
21,728 votes
$4.13
$0.75 / $3.75 per M tokens
1,573Independently tested
21 results
Claude Sonnet 5
Anthropic
156
153–159
11th of 14
43,340 votes
$11
$2 / $10 per M tokens
1,500Independently tested
21 results
GPT-5.6 Luna
OpenAI
156
154–159
13th of 14
35,823 votes
$1.24
$0.2 / $1.2 per M tokens
1,575Independently tested
23 results
DeepSeek V4 Pro
DeepSeek · open weights
155
154–157
9th of 14
9,857 votes
$1.48
$0.435 / $0.87 per M tokens
1,500Independently tested
17 results
DeepSeek V4.1 Flash
DeepSeek · open weights
155
149–158
7th of 14
6,630 votes
$0.72
$0.15 / $0.6 per M tokens
1,500Independently tested
6 results
Gemini 3.1 Pro
Google DeepMind
155
153–158
4th of 14
119,196 votes
$12
$2 / $12 per M tokens
1,573Independently tested
34 results
Gemini 3.6 Flash
Google DeepMind
154
153–156
5th of 14
33,739 votes
$4.13
$0.75 / $3.75 per M tokens
1,573Independently tested
20 results
Gemini 3.5 Flash-Lite
Google DeepMind
145
142–147
12th of 14
33,501 votes
$2.35
$0.3 / $2.5 per M tokens
1,573Independently tested
14 results
Claude Haiku 4.5
Anthropic
142
139–144
14th of 14
140,791 votes
$5.50
$1 / $5 per M tokens
300Independently tested
17 results
Claude Opus 5.5
Anthropic
–1st of 14
2,307 votes
$22
$4 / $20 per M tokens
1,500Early results
2 results
GPT-6 Luna
OpenAI
––$0.55
$0.1 / $0.5 per M tokens
1,575Early results
1 result
GPT-6 Sol
OpenAI
––$11
$2 / $10 per M tokens
1,575Early results
2 results

Sources: Epoch AI Benchmarking Hub (CC BY 4.0) · LMArena text leaderboard (CC BY 4.0) · models.dev (MIT) · model makers' announcements. How it works

Intelligence Snacks newsletter

The big AI ideas each week, in your inbox