PodcastIntelligence Snacks · Weekly conversations about AI, software and business, from the team behind Normie Mode.Listen →
Best AI for…

Best AI for research

Finding information, checking facts and pulling it all together.

Best overall, and best for $25 a month or lessToo close to call
  • GPT-6 AstraGet it with ChatGPT Plus, $20 a month (only in ChatGPT's Work mode and Codex, not normal chat)
  • GPT-5.6 SolGet it with ChatGPT Plus, $20 a month

1st and 2nd of 11 models for this task

Best freeToo close to call

3rd and 4th of 11 models for this task

Keep in mind Facts and expert-question tests help, but nothing yet tests research reports or citations.

Based on tests of related skills. The picks follow a fixed rule using the charts below and each app's plans.

Already paying for one? See what each plan gives you here
App and planPriceBest model on it for thisPosition
ChatGPT FreefreeGPT-5.6 Luna9th of 11
ChatGPT Go$8 a monthGPT-5.6 Luna9th of 11
ChatGPT Plus$20 a monthGPT-6 Astra
Only in ChatGPT's Work mode and Codex, not normal chat
1st of 11
ChatGPT Pro$100 a monthGPT-6 Astra
Called GPT-6 Pro in the app
1st of 11
Claude FreefreeClaude Sonnet 510th of 11
Claude Pro$20 a monthClaude Sonnet 5
Its main model, Claude Opus 5.5, is too new to score yet.
10th of 11
Claude Max$100 a monthClaude Fable 5.1
For up to half your weekly allowance
Its main model, Claude Opus 5.5, is too new to score yet.
5th of 11
Gemini FreefreeGemini 3.6 Flash3rd of 11
Google AI Plus$4.99 a monthGemini 3.6 Flash3rd of 11
Google AI Pro$19.99 a monthGemini 3.8 Flash6th of 11
Google AI Ultra$99.99 a monthGemini 3.8 Flash6th of 11
Kimi FreefreeKimi K34th of 11
Kimi Moderato$19 a monthKimi K34th of 11

Norm's guide

What I'd tell a friend who asked.

What AI is good at

  • Summarising long articles and reports.
  • Explaining a topic you're new to.
  • Comparing a few options side by side.

Where it trips up

  • Inventing sources, quotes or statistics.
  • Out-of-date information.
  • Presenting a guess as if it were a fact.

How to get a good result

  • Ask for sources, then open them yourself.
  • Use a model that can search the web for anything recent.
  • Double-check any number you plan to repeat.

How the models compare

Every model we could score for research, combining the tests below.

Overall for research

Position on a combined score from 3 tests. The longer the bar, the further ahead · Higher is better

On the tests we have, GPT-6 Astra comes out on top, followed by GPT-5.6 Sol and Gemini 3.6 Flash.

GPT-6 Astra · OpenAI1st
GPT-5.6 Sol · OpenAI2nd
Gemini 3.6 Flash · Google DeepMind3rd
Kimi K3 · Moonshot AI4th
Claude Fable 5.1 · Anthropic5th
Gemini 3.8 Flash · Google DeepMind6th
Gemini 3.1 Pro · Google DeepMind7th
DeepSeek V4 Pro · DeepSeek8th
GPT-5.6 Luna · OpenAI9th
Claude Sonnet 5 · Anthropic10th
Claude Haiku 4.5 · Anthropic11th

The evidence

Answering short factual questions correctly

Tests a related skill · Short questions that each have one checkable answer. · Higher is better

GPT-6 Astra and Gemini 3.1 Pro are neck and neck at the top, followed by Claude Fable 5.1 and Gemini 3.8 Flash; GPT-6 Astra costs about 4 times as much.

GPT-6 Astra · OpenAI76%
Gemini 3.1 Pro · Google DeepMind74%
Claude Fable 5.1 · Anthropic71%
Gemini 3.8 Flash · Google DeepMind70%
GPT-5.6 Sol · OpenAI70%
Gemini 3.6 Flash · Google DeepMind66%
DeepSeek V4 Pro · DeepSeek53%
Kimi K3 · Moonshot AI51%
GPT-5.6 Luna · OpenAI41%
Claude Sonnet 5 · Anthropic34%
Claude Haiku 4.5 · Anthropic13%
Source: Google DeepMind via Epoch AI · CC BY 4.0

Humanity's Last Exam

Tests a related skill · Very hard questions written by experts across many subjects. · Higher is better

GPT-6 Astra leads, followed by Claude Fable 5.1 and Gemini 3.1 Pro.

GPT-6 Astra · OpenAI55%
Claude Fable 5.1 · Anthropic47%
Gemini 3.1 Pro · Google DeepMind46%
Gemini 3.8 Flash · Google DeepMind45%

People's votes on expert questions

Tests a related skill · People compared two anonymous answers to questions needing expert knowledge. · Higher is better

Claude Opus 5.5 and Gemini 3.8 Flash are neck and neck at the top, followed by Claude Fable 5.1 and Kimi K3; Claude Opus 5.5 costs about 5 times as much.

Claude Opus 5.5 · Anthropic1st
Gemini 3.8 Flash · Google DeepMind2nd
Claude Fable 5.1 · Anthropic3rd
Kimi K3 · Moonshot AI4th
GPT-5.6 Sol · OpenAI5th
DeepSeek V4.1 Flash · DeepSeek6th
Claude Sonnet 5 · Anthropic7th
Gemini 3.6 Flash · Google DeepMind8th
GPT-6 Astra · OpenAI9th
Gemini 3.1 Pro · Google DeepMind10th
GPT-5.6 Luna · OpenAI11th
DeepSeek V4 Pro · DeepSeek12th
Claude Haiku 4.5 · Anthropic13th
Gemini 3.5 Flash-Lite · Google DeepMind14th
Source: LMArena · votes as of 25 Sept 2026

If you build with it

What the makers charge developers who use the models directly. Using an app? The plans above are what you pay.

Cost per 1,000 typical requests

List price, about 1,500 words in and 500 out per request · Lower is better

GPT-6 Luna is cheapest, followed by DeepSeek V4.1 Flash and GPT-5.6 Luna.

GPT-6 Luna · OpenAI$0.55
DeepSeek V4.1 Flash · DeepSeek$0.72
GPT-5.6 Luna · OpenAI$1.24
DeepSeek V4 Pro · DeepSeek$1.48
Gemini 3.5 Flash-Lite · Google DeepMind$2.35
Gemini 3.8 Flash · Google DeepMind$4.13
Gemini 3.6 Flash · Google DeepMind$4.13
Claude Haiku 4.5 · Anthropic$5.50
Claude Sonnet 5 · Anthropic$11
GPT-6 Sol · OpenAI$11
Gemini 3.1 Pro · Google DeepMind$12
Kimi K3 · Moonshot AI$17
GPT-5.6 Sol · OpenAI$22
Claude Opus 5.5 · Anthropic$22
GPT-6 Astra · OpenAI$55
Claude Fable 5.1 · Anthropic$55
Source: models.dev · MIT
Intelligence Snacks newsletter

The big AI ideas each week, in your inbox