PodcastIntelligence Snacks · Weekly conversations about AI, software and business, from the team behind Normie Mode.Listen →
Tests explained

The tests behind our rankings

What each test checks, in one sentence, and which tasks it helps us rank.

Tested by us

Tests we design and run ourselves, with every question and every model's answer published.

Tests of real work

Set tasks with right and wrong answers, run by independent testers.

Professional tasks in banking, consulting and law

Tasks written by experienced professionals that take a person about two hours, using documents, spreadsheets, email and slides.

Mercor · 8 models · 7 tasks

Graduate-level science questions

Biology, physics and chemistry questions written so they can't be answered by searching the web.

NYU and others · 11 models · 5 tasks

Pulling information out of documents into tables

Reading documents and filling in a table with the right figures, including cases that need reasoning.

DTBench authors · 13 models · 5 tasks

Answering short factual questions correctly

Short questions that each have one checkable answer.

Google DeepMind · 11 models · 2 tasks

Competition maths

Problems in the style of a top US high-school maths competition.

Epoch AI · 12 models · 2 tasks

Humanity's Last Exam

Very hard questions written by experts across many subjects.

Center for AI Safety, Scale AI · 4 models · 2 tasks

Building features in real software projects

113 original tasks across 91 real code projects in five languages, checked with hand-written tests.

Datacurve · 8 models · 1 task

Getting jobs done on a computer's command line

Practical tasks completed by typing commands, like setting up software or processing files.

Terminal-Bench team · 2 models · 1 task

Real paid freelance projects

Projects originally done by freelancers for pay, scored on whether the AI's result would be acceptable to the client.

Scale AI, Center for AI Safety · 2 models · 1 task

Research-level maths

New, unpublished maths problems that take specialists hours or days.

Epoch AI · 11 models · 1 task

Would a real maintainer accept its code?

Coding tasks in 36 major open-source projects, judged on whether the code is good enough to merge, not just whether it runs.

Cognition · 9 models · 1 task

People's votes

People ask a question, see two anonymous answers and pick the better one. It shows which answers people prefer, not whether they were right.

People's votes on writing and language

People compared two anonymous answers to writing, editing and language questions.

LMArena · 14 models · 9 tasks

People's votes on business and finance questions

People compared two anonymous answers to business, management and finance questions.

LMArena · 14 models · 6 tasks

People's votes on expert questions

People compared two anonymous answers to questions needing expert knowledge.

LMArena · 14 models · 6 tasks

People's votes on following instructions

People compared two anonymous answers to requests with specific instructions.

LMArena · 14 models · 6 tasks

People's votes on creative writing

People compared two anonymous pieces of creative writing.

LMArena · 14 models · 4 tasks

People's votes on back-and-forth conversations

People compared two anonymous answers in conversations with several back-and-forth turns.

LMArena · 14 models · 3 tasks

People's votes on health questions

People compared two anonymous answers to medical and healthcare questions.

LMArena · 13 models · 3 tasks

People's votes on long, detailed requests

People compared two anonymous answers to long requests with lots of detail to take in.

LMArena · 14 models · 3 tasks

People's votes on science questions

People compared two anonymous answers to questions in the life, physical and social sciences.

LMArena · 14 models · 3 tasks

People's votes on maths questions

People compared two anonymous answers to maths questions.

LMArena · 13 models · 2 tasks

People's votes on questions in other languages

People compared two anonymous answers to questions asked in languages other than English.

LMArena · 14 models · 2 tasks

People's votes on coding questions

People compared two anonymous answers to their coding questions and picked the better one.

LMArena · 14 models · 1 task

People's votes on legal questions

People compared two anonymous answers to legal and government questions.

LMArena · 14 models · 1 task

People's votes overall

People compared two anonymous answers to all kinds of questions and picked the better one.

LMArena · 14 models · 1 task

Intelligence Snacks newsletter

The big AI ideas each week, in your inbox