Best AI for coding and building apps
Building apps and websites, fixing bugs and automating tasks.
1st and 2nd of 10 models for this task
Get it with ChatGPT Plus, $20 a month (only in ChatGPT's Work mode and Codex, not normal chat)
2nd of 10 models for this task
Based on tests of this task. The picks follow a fixed rule using the charts below and each app's plans.
Norm's guide
What I'd tell a friend who asked.
What AI is good at
- Writing small scripts and tools from a plain-English description.
- Explaining code you didn't write.
- Finding and fixing bugs.
Where it trips up
- Large projects with lots of moving parts.
- Code that runs but isn't quite what you asked for.
- Security details, like passwords and permissions.
How to get a good result
- Describe what you want in small steps.
- Test the result before you rely on it.
- Ask it to explain anything you don't understand.
How the models compare
Every model we could score for coding and building apps, combining the tests below.
Overall for coding and building apps
Position on a combined score from 3 tests. The longer the bar, the further ahead · Higher is better
On the tests we have, Claude Fable 5.1 comes out on top, followed by GPT-6 Astra and GPT-5.6 Sol.
The evidence
Would a real maintainer accept its code?
Tests this task · Coding tasks in 36 major open-source projects, judged on whether the code is good enough to merge, not just whether it runs. · Higher is better
GPT-6 Astra leads, followed by Claude Fable 5.1 and GPT-5.6 Sol.
Building features in real software projects
Tests this task · 113 original tasks across 91 real code projects in five languages, checked with hand-written tests. · Higher is better
GPT-6 Astra leads, followed by Gemini 3.8 Flash and GPT-5.6 Sol.
People's votes on coding questions
Tests this task · People compared two anonymous answers to their coding questions and picked the better one. · Higher is better
Claude Opus 5.5 and Claude Fable 5.1 are neck and neck at the top, followed by Gemini 3.8 Flash and DeepSeek V4.1 Flash; Claude Fable 5.1 costs about 3 times as much.
Also relevant, but too few of our models tested to compare yet: Getting jobs done on a computer's command line (2 models).
Also relevant, but no results for the models we compare yet: Fixing reported bugs in real software.
If you build with it
What the makers charge developers who use the models directly. Using an app? The plans above are what you pay.
Cost per 1,000 typical requests
List price, about 1,500 words in and 500 out per request · Lower is better
GPT-6 Luna is cheapest, followed by DeepSeek V4.1 Flash and GPT-5.6 Luna.