Coding and building apps
Building apps and websites, fixing bugs and automating tasks.
Claude Fable 5.11stGPT-6 Astra2ndGPT-5.6 Sol3rd- Up to $25
- GPT-6 Astra (ChatGPT Plus)
- Free
- Kimi K3 (Kimi Free)
Tests this task · 3 tests
Pick your task to see which AI does it best, where to get it, and the evidence behind our pick.
Updated 26 Sept 2026
The everyday jobs people search for, most popular first, with the models that do each best.
Building apps and websites, fixing bugs and automating tasks.
Claude Fable 5.11stGPT-6 Astra2ndGPT-5.6 Sol3rdTests this task · 3 tests
Emails, articles, reports, stories and editing.
Claude Opus 5.51stClaude Fable 5.12ndGemini 3.8 Flash3rdTests this task · 3 tests
Homework, revision, exam practice and learning new topics.
Gemini 3.8 Flash1stClaude Fable 5.12ndKimi K33rdTests a related skill · 4 tests
CVs, cover letters, job applications and interview practice.
Claude Opus 5.51stClaude Fable 5.12ndGemini 3.8 Flash3rdTests a related skill · 2 tests
Planning, strategy, brainstorming and everyday office work.
Claude Fable 5.11stGemini 3.8 Flash2ndGPT-6 Astra3rdTests this task · 2 tests
Finding information, checking facts and pulling it all together.
GPT-6 Astra1stGPT-5.6 Sol2ndGemini 3.6 Flash3rdTests a related skill · 3 tests
Understanding symptoms, conditions and medical information.
Claude Opus 5.51stClaude Fable 5.12ndGemini 3.8 Flash3rdTests this task · 2 tests
Reading PDFs and documents, pulling out information and taking notes.
Claude Fable 5.11stGemini 3.8 Flash2ndGemini 3.1 Pro3rdTests this task · 2 tests
Accounting, financial analysis, budgets and planning.
Claude Fable 5.11stGemini 3.8 Flash2ndGPT-6 Astra3rdTests a related skill · 2 tests
How the best model on each plan does across the 15 everyday tasks above.
| App and plan | Price | In the top three for |
|---|---|---|
| ChatGPT Free | Free | 0 of 15 tasks |
| ChatGPT Go | $8/month | 0 of 15 tasks |
| ChatGPT Plus | $20/month | 7 of 15 tasks coding and building apps, business, research, accounting and finance… |
| ChatGPT Pro | $100/month | 7 of 15 tasks coding and building apps, business, research, accounting and finance… |
| Claude Free | Free | 0 of 15 tasks |
| Claude Pro | $20/month | 5 of 15 tasks writing, cvs, job applications and interviews, health questions, planning and personal admin… |
| Claude Max | $100/month | 14 of 15 tasks coding and building apps, writing, studying and homework, cvs, job applications and interviews… |
| Gemini Free | Free | 2 of 15 tasks research, documents, pdfs and notes |
| Google AI Plus | $4.99/month | 2 of 15 tasks research, documents, pdfs and notes |
| Google AI Pro | $19.99/month | 13 of 15 tasks writing, studying and homework, cvs, job applications and interviews, business… |
| Google AI Ultra | $99.99/month | 13 of 15 tasks writing, studying and homework, cvs, job applications and interviews, business… |
| Kimi Free | Free | 2 of 15 tasks studying and homework, maths and science |
| Kimi Moderato | $19/month | 2 of 15 tasks studying and homework, maths and science |
The models behind the apps, side by side: how capable, how much people like them and what they cost developers.
GPT-6 Astra and Claude Fable 5.1 are the most capable overall, followed by GPT-5.6 Sol and Kimi K3.
Claude Opus 5.5 and Claude Fable 5.1 are preferred most in blind votes, followed by Gemini 3.8 Flash and Gemini 3.1 Pro.
GPT-6 Luna and DeepSeek V4.1 Flash cost the least to use, followed by GPT-5.6 Luna and DeepSeek V4 Pro.
GPT-6 Astra and GPT-5.6 Sol can take in the most text at once, followed by GPT-5.6 Luna and GPT-6 Luna.
Claude Opus 5.5, GPT-6 Luna and GPT-6 Sol came out in the last few days and have only 1–2 independent results so far.
Position on a score combining dozens of tests · Higher is better
Position in blind votes on LMArena · Higher is better
Per 1,000 typical requests · Lower is better
One score that combines dozens of independent tests. The thin line on each bar is the likely range.
Position on Epoch AI's score, which combines dozens of independent tests. The thin line shows the likely range · Higher is better
If two models' ranges overlap, treat them as roughly equal.
Up and to the left is better: more capable for less money.
Epoch capability score against list price per 1,000 typical requests (log scale)
People compare two anonymous answers and vote for the better one.
Position among the models here, from votes as of 25 Sept 2026 · Higher is better
New models appear here after they collect enough votes, usually within a couple of weeks.
What you'd pay the maker directly. Chat apps with monthly plans price differently.
List price, assuming about 1,500 words in, 500 out per request. Models that think longer cost more in practice · Lower is better
How much text it can take in at once, in pages of about 500 words.
Based on each model's context window · Higher is better
What the makers say, until independent results arrive.
The maker says“It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.”
Independent results so far: 2
The maker says“At high effort, GPT-6 Luna improves on its predecessor by 5.4 percentage points at 58% lower cost per task.”
Independent results so far: 1
The maker says“On AutomationBench, GPT-6 Sol at xhigh effort outperforms Claude Opus 5 at max effort at just 9% of Opus 5's cost per task.”
Independent results so far: 2
Everything above in one table.
| Model | Capability | People's votes | Cost / 1,000 requests | Pages at once | Evidence |
|---|---|---|---|---|---|
| GPT-6 Astra OpenAI | 167 163–172 | 10th of 14 7,185 votes | $55 $10 / $50 per M tokens | 1,575 | Independently tested 27 results |
| Claude Fable 5.1 Anthropic | 165 162–170 | 2nd of 14 9,942 votes | $55 $10 / $50 per M tokens | 1,500 | Independently tested 25 results |
| GPT-5.6 Sol OpenAI | 162 160–166 | 8th of 14 34,258 votes | $22 $4 / $20 per M tokens | 1,575 | Independently tested 29 results |
| Kimi K3 Moonshot AI · open weights | 158 155–161 | 6th of 14 26,400 votes | $17 $3 / $15 per M tokens | 1,573 | Independently tested 24 results |
| Gemini 3.8 Flash Google DeepMind | 157 155–161 | 3rd of 14 21,728 votes | $4.13 $0.75 / $3.75 per M tokens | 1,573 | Independently tested 21 results |
| Claude Sonnet 5 Anthropic | 156 153–159 | 11th of 14 43,340 votes | $11 $2 / $10 per M tokens | 1,500 | Independently tested 21 results |
| GPT-5.6 Luna OpenAI | 156 154–159 | 13th of 14 35,823 votes | $1.24 $0.2 / $1.2 per M tokens | 1,575 | Independently tested 23 results |
| DeepSeek V4 Pro DeepSeek · open weights | 155 154–157 | 9th of 14 9,857 votes | $1.48 $0.435 / $0.87 per M tokens | 1,500 | Independently tested 17 results |
| DeepSeek V4.1 Flash DeepSeek · open weights | 155 149–158 | 7th of 14 6,630 votes | $0.72 $0.15 / $0.6 per M tokens | 1,500 | Independently tested 6 results |
| Gemini 3.1 Pro Google DeepMind | 155 153–158 | 4th of 14 119,196 votes | $12 $2 / $12 per M tokens | 1,573 | Independently tested 34 results |
| Gemini 3.6 Flash Google DeepMind | 154 153–156 | 5th of 14 33,739 votes | $4.13 $0.75 / $3.75 per M tokens | 1,573 | Independently tested 20 results |
| Gemini 3.5 Flash-Lite Google DeepMind | 145 142–147 | 12th of 14 33,501 votes | $2.35 $0.3 / $2.5 per M tokens | 1,573 | Independently tested 14 results |
| Claude Haiku 4.5 Anthropic | 142 139–144 | 14th of 14 140,791 votes | $5.50 $1 / $5 per M tokens | 300 | Independently tested 17 results |
| Claude Opus 5.5 Anthropic | – | 1st of 14 2,307 votes | $22 $4 / $20 per M tokens | 1,500 | Early results 2 results |
| GPT-6 Luna OpenAI | – | – | $0.55 $0.1 / $0.5 per M tokens | 1,575 | Early results 1 result |
| GPT-6 Sol OpenAI | – | – | $11 $2 / $10 per M tokens | 1,575 | Early results 2 results |
Sources: Epoch AI Benchmarking Hub (CC BY 4.0) · LMArena text leaderboard (CC BY 4.0) · models.dev (MIT) · model makers' announcements. How it works