PodcastIntelligence Snacks · Weekly conversations about AI, software and business, from the team behind Normie Mode.Listen →
Best AI for… / coding and building apps

Best AI for building apps

Making phone and desktop apps, from idea to working version.

Best overallToo close to call
  • Claude Fable 5.1Get it with Claude Max, $100 a month (for up to half your weekly allowance)
  • GPT-6 AstraGet it with ChatGPT Plus, $20 a month (only in ChatGPT's Work mode and Codex, not normal chat)

1st and 2nd of 10 models for this task

Best for $25 a month or less

GPT-6 Astra

Get it with ChatGPT Plus, $20 a month (only in ChatGPT's Work mode and Codex, not normal chat)

2nd of 10 models for this task

Best free

Kimi K3

Get it with Kimi Free

4th of 10 models for this task

Based on the tests for coding and building apps. The picks follow a fixed rule using the charts below and each app's plans.

Already paying for one? See what each plan gives you here
App and planPriceBest model on it for thisPosition
ChatGPT FreefreeGPT-5.6 Luna7th of 10
ChatGPT Go$8 a monthGPT-5.6 Luna7th of 10
ChatGPT Plus$20 a monthGPT-6 Astra
Only in ChatGPT's Work mode and Codex, not normal chat
2nd of 10
ChatGPT Pro$100 a monthGPT-6 Astra
Called GPT-6 Pro in the app
2nd of 10
Claude FreefreeClaude Sonnet 56th of 10
Claude Pro$20 a monthClaude Sonnet 5
Its main model, Claude Opus 5.5, is too new to score yet.
6th of 10
Claude Max$100 a monthClaude Fable 5.1
For up to half your weekly allowance
Its main model, Claude Opus 5.5, is too new to score yet.
1st of 10
Gemini FreefreeGemini 3.6 Flash8th of 10
Google AI Plus$4.99 a monthGemini 3.6 Flash8th of 10
Google AI Pro$19.99 a monthGemini 3.8 Flash5th of 10
Google AI Ultra$99.99 a monthGemini 3.8 Flash5th of 10
Kimi FreefreeKimi K34th of 10
Kimi Moderato$19 a monthKimi K34th of 10

How the models compare

Every model we could score for building apps, combining the tests below.

Overall for building apps

Position on a combined score from 3 tests. The longer the bar, the further ahead · Higher is better

On the tests we have, Claude Fable 5.1 comes out on top, followed by GPT-6 Astra and GPT-5.6 Sol.

Claude Fable 5.1 · Anthropic1st
GPT-6 Astra · OpenAI2nd
GPT-5.6 Sol · OpenAI3rd
Kimi K3 · Moonshot AI4th
Gemini 3.8 Flash · Google DeepMind5th
Claude Sonnet 5 · Anthropic6th
GPT-5.6 Luna · OpenAI7th
Gemini 3.6 Flash · Google DeepMind8th
Gemini 3.1 Pro · Google DeepMind9th
DeepSeek V4 Pro · DeepSeek10th

The evidence

Would a real maintainer accept its code?

Tests a related skill · Coding tasks in 36 major open-source projects, judged on whether the code is good enough to merge, not just whether it runs. · Higher is better

GPT-6 Astra leads, followed by Claude Fable 5.1 and GPT-5.6 Sol.

GPT-6 Astra · OpenAI53%
Claude Fable 5.1 · Anthropic51%
GPT-5.6 Sol · OpenAI47%
Kimi K3 · Moonshot AI44%
Claude Sonnet 5 · Anthropic43%
Gemini 3.8 Flash · Google DeepMind41%
GPT-5.6 Luna · OpenAI40%
Gemini 3.6 Flash · Google DeepMind34%
DeepSeek V4 Pro · DeepSeek29%
Source: Cognition via Epoch AI · CC BY 4.0

Building features in real software projects

Tests a related skill · 113 original tasks across 91 real code projects in five languages, checked with hand-written tests. · Higher is better

GPT-6 Astra leads, followed by Gemini 3.8 Flash and GPT-5.6 Sol.

GPT-6 Astra · OpenAI74%
Gemini 3.8 Flash · Google DeepMind74%
GPT-5.6 Sol · OpenAI73%
Kimi K3 · Moonshot AI69%
GPT-5.6 Luna · OpenAI67%
Claude Sonnet 5 · Anthropic54%
Gemini 3.6 Flash · Google DeepMind47%
Gemini 3.1 Pro · Google DeepMind12%
Source: Datacurve via Epoch AI · CC BY 4.0

People's votes on coding questions

Tests a related skill · People compared two anonymous answers to their coding questions and picked the better one. · Higher is better

Claude Opus 5.5 and Claude Fable 5.1 are neck and neck at the top, followed by Gemini 3.8 Flash and DeepSeek V4.1 Flash; Claude Fable 5.1 costs about 3 times as much.

Claude Opus 5.5 · Anthropic1st
Claude Fable 5.1 · Anthropic2nd
Gemini 3.8 Flash · Google DeepMind3rd
DeepSeek V4.1 Flash · DeepSeek4th
Kimi K3 · Moonshot AI5th
GPT-5.6 Sol · OpenAI6th
Gemini 3.6 Flash · Google DeepMind7th
GPT-6 Astra · OpenAI8th
Claude Sonnet 5 · Anthropic9th
Gemini 3.1 Pro · Google DeepMind10th
DeepSeek V4 Pro · DeepSeek11th
GPT-5.6 Luna · OpenAI12th
Gemini 3.5 Flash-Lite · Google DeepMind13th
Claude Haiku 4.5 · Anthropic14th
Source: LMArena · votes as of 25 Sept 2026

Also relevant, but too few of our models tested to compare yet: Getting jobs done on a computer's command line (2 models).

Also relevant, but no results for the models we compare yet: Fixing reported bugs in real software.

If you build with it

What the makers charge developers who use the models directly. Using an app? The plans above are what you pay.

Cost per 1,000 typical requests

List price, about 1,500 words in and 500 out per request · Lower is better

GPT-6 Luna is cheapest, followed by DeepSeek V4.1 Flash and GPT-5.6 Luna.

GPT-6 Luna · OpenAI$0.55
DeepSeek V4.1 Flash · DeepSeek$0.72
GPT-5.6 Luna · OpenAI$1.24
DeepSeek V4 Pro · DeepSeek$1.48
Gemini 3.5 Flash-Lite · Google DeepMind$2.35
Gemini 3.8 Flash · Google DeepMind$4.13
Gemini 3.6 Flash · Google DeepMind$4.13
Claude Haiku 4.5 · Anthropic$5.50
Claude Sonnet 5 · Anthropic$11
GPT-6 Sol · OpenAI$11
Gemini 3.1 Pro · Google DeepMind$12
Kimi K3 · Moonshot AI$17
GPT-5.6 Sol · OpenAI$22
Claude Opus 5.5 · Anthropic$22
GPT-6 Astra · OpenAI$55
Claude Fable 5.1 · Anthropic$55
Source: models.dev · MIT
Intelligence Snacks newsletter

The big AI ideas each week, in your inbox