Best AI for research
Finding information, checking facts and pulling it all together.
1st and 2nd of 11 models for this task
3rd and 4th of 11 models for this task
Keep in mind Facts and expert-question tests help, but nothing yet tests research reports or citations.
Based on tests of related skills. The picks follow a fixed rule using the charts below and each app's plans.
Norm's guide
What I'd tell a friend who asked.
What AI is good at
- Summarising long articles and reports.
- Explaining a topic you're new to.
- Comparing a few options side by side.
Where it trips up
- Inventing sources, quotes or statistics.
- Out-of-date information.
- Presenting a guess as if it were a fact.
How to get a good result
- Ask for sources, then open them yourself.
- Use a model that can search the web for anything recent.
- Double-check any number you plan to repeat.
How the models compare
Every model we could score for research, combining the tests below.
Overall for research
Position on a combined score from 3 tests. The longer the bar, the further ahead · Higher is better
On the tests we have, GPT-6 Astra comes out on top, followed by GPT-5.6 Sol and Gemini 3.6 Flash.
The evidence
Answering short factual questions correctly
Tests a related skill · Short questions that each have one checkable answer. · Higher is better
GPT-6 Astra and Gemini 3.1 Pro are neck and neck at the top, followed by Claude Fable 5.1 and Gemini 3.8 Flash; GPT-6 Astra costs about 4 times as much.
Humanity's Last Exam
Tests a related skill · Very hard questions written by experts across many subjects. · Higher is better
GPT-6 Astra leads, followed by Claude Fable 5.1 and Gemini 3.1 Pro.
People's votes on expert questions
Tests a related skill · People compared two anonymous answers to questions needing expert knowledge. · Higher is better
Claude Opus 5.5 and Gemini 3.8 Flash are neck and neck at the top, followed by Claude Fable 5.1 and Kimi K3; Claude Opus 5.5 costs about 5 times as much.
If you build with it
What the makers charge developers who use the models directly. Using an app? The plans above are what you pay.
Cost per 1,000 typical requests
List price, about 1,500 words in and 500 out per request · Lower is better
GPT-6 Luna is cheapest, followed by DeepSeek V4.1 Flash and GPT-5.6 Luna.