How it works
We turn independent tests and people's votes into a pick for your task, then tell you which app and plan gets it. Here's each step, using the site's own data.
16 from Anthropic, OpenAI, Google, DeepSeek and Moonshot AI.
26, each explained on Tests explained.
46 task pages, named after what people search for.
26 Sept 2026. Plans checked 26 Sept 2026.
From test to pick
Four steps, shown for coding and building apps, the task people search for most.
- 1
Gather the evidence
Tests and votes that measure the task, each labelled by how closely it matches.
3 tests for coding and building apps
- 2
Score each model
Combine its place on every test into one task score.
Claude Fable 5.1 comes 1st of 10
- 3
Make the picks
Best overall, best for $25 a month or less, and best free.
3 picks
- 4
Match to a plan
Find the cheapest app and plan that includes each pick.
Claude Max, $100/month
Where the evidence comes from
Our own tests, plus free, public sources whose licences let us show their data. Every chart names its source underneath, and our own results carry the navy Tested by us tag.
Tested by usOurs
Our own tests of everyday work, with every question and every model's answer published
- We use
- 1 test
- Licence
- Ours
Epoch AI
Results from independent tests of real work, like fixing code or answering expert questions
- We use
- 11 tests
- Licence
- CC BY 4.0
- As of
- 26 Sept 2026
LMArena
People comparing two anonymous answers and picking the better one
- We use
- 14 vote categories
- Licence
- CC BY 4.0
- As of
- 25 Sept 2026
models.dev
Prices, release dates and how much each model can read
- We use
- 16 models
- Licence
- MIT
- As of
- 26 Sept 2026
Makers' own pages
Which plans include which model, and what makers claim at launch (always labelled)
- We use
- 4 apps, 13 plans
- Licence
- Cited
- As of
- 26 Sept 2026
How close a test is to your task
Hardly any test measures an everyday task exactly, so every test on a task page carries one of three labels.
The test is this kind of work.
17 task pages rest mainly on this
It tests overall ability. Used only when nothing closer exists.
0 task pages rest mainly on this
A more specific task, like vibe coding, borrows its broader task's tests when it has none of its own, labelled as related at most.
The task score
On each test, the lowest of our models is placed at 0 and the highest at 100. A model's task score is the average of its places. Here's how the top of coding and building apps is worked out:
| Model | Would a real maintainer accept its code? Tests this task | Building features in real software projects Tests this task | People's votes on coding questions Tests this task | Task score |
|---|---|---|---|---|
| Claude Fable 5.1 1st | 90 | – | 75 | 83 |
| GPT-6 Astra 2nd | 100 | 100 | 38 | 79 |
| GPT-5.6 Sol 3rd | 77 | 98 | 43 | 72 |
| Kimi K3 4th | 63 | 91 | 61 | 72 |
- It's relative. It shows who leads among the models we track, not how good they are in absolute terms. Last place isn't useless.
- A test needs at least 3 of our models before it counts, so two results can't decide a ranking.
- A model needs results on at least half the tests, including at least one test of real work when the task has one. People's votes alone can't put a model on the chart.
- A dash means the model hasn't taken that test; its score averages the tests it has.
4 more models have some results for coding and building apps but not enough to score yet.
The picks
Every task page opens with picks made by a fixed rule, so they always agree with the charts:
- Best overall: the highest task score, and the cheapest plan that includes it.
- Best for $25 a month or less: the highest-scoring model on a plan at that price or below.
- Best free: the highest-scoring model on a free plan.
- Within 5 points: it's too close to call, so we name both, higher-ranked first, each with its plan and price.
- Extra cost doesn't count. A model you pay for on top of a plan isn't treated as included.
Here's the result for coding and building apps:
1st and 2nd of 10 models for this task
Get it with ChatGPT Plus, $20 a month (only in ChatGPT's Work mode and Codex, not normal chat)
2nd of 10 models for this task
Based on tests of this task. The picks follow a fixed rule using the charts below and each app's plans.
Too close to call
Results come with a likely range, drawn as a thin line on each bar. When the top two lines overlap, we say it's too close to call rather than naming a winner.
People's votes on business and finance questions
An example from the site · Higher is better
New models
Independent testers take one to three weeks to catch up with a launch, so each model shows where its evidence stands. A new version always starts from nothing; we never assume it performs like the one before.
8 or more, or a place on Epoch AI's combined score.
13 models right now
Fewer than 8 independent results.
3 models right now
No independent results yet, only what the maker says.
0 models right now
Apps and plans
ChatGPT, Claude and Gemini are apps; each runs one or more models, and free plans often run an older one. We compare the models, then map each to the plans that include it.
| App | Plans we cover | Checked |
|---|---|---|
| ChatGPT OpenAI | Free (Free), Go ($8/month), Plus ($20/month), Pro ($100/month) | 26 Sept 2026 |
| Claude Anthropic | Free (Free), Pro ($20/month), Max ($100/month) | 26 Sept 2026 |
| Gemini Google | Free (Free), Google AI Plus ($4.99/month), Google AI Pro ($19.99/month), Google AI Ultra ($99.99/month) | 26 Sept 2026 |
| Kimi Moonshot AI | Free (Free), Moderato ($19/month) | 26 Sept 2026 |
- We check plans against the makers' own pages at least every 30 days. Where a maker's page can't be read, we use press coverage and say so.
- An older model that a plan still uses stays on our charts, because it's what people on that plan actually get.
- Plans change often, so check the price before you buy.
Cost charts show what developers pay to use a model directly, for 1,000 typical requests of about 1,500 words in, 500 out. If you use an app, the monthly plan is what you pay.
Who writes what

Norm, our guide
Anything Norm says about the data, like who leads or what the top models cost, is filled in from the same numbers as the charts, so it can't disagree with them. His advice is written by a person and only appears once it's been reviewed.
What the makers say
Launch claims appear on model pages, always labelled as the maker's own, and never count towards a score.
What it can't tell you
- How good AI is at your exact job. Tests are close stand-ins, and the label on each chart tells you how close.
- Whether one answer is right. Every model makes mistakes; Norm's guide on each task says what to check.
- How a model does inside a specific app feature, like a document upload or a coding tool. We compare the models, not the apps.