The tests behind our rankings
What each test checks, in one sentence, and which tasks it helps us rank.
Tested by us
Tests we design and run ourselves, with every question and every model's answer published.
Tests of real work
Set tasks with right and wrong answers, run by independent testers.
Professional tasks in banking, consulting and law
Tasks written by experienced professionals that take a person about two hours, using documents, spreadsheets, email and slides.
Mercor · 8 models · 7 tasks
Graduate-level science questions
Biology, physics and chemistry questions written so they can't be answered by searching the web.
NYU and others · 11 models · 5 tasks
Pulling information out of documents into tables
Reading documents and filling in a table with the right figures, including cases that need reasoning.
DTBench authors · 13 models · 5 tasks
Answering short factual questions correctly
Short questions that each have one checkable answer.
Google DeepMind · 11 models · 2 tasks
Competition maths
Problems in the style of a top US high-school maths competition.
Epoch AI · 12 models · 2 tasks
Humanity's Last Exam
Very hard questions written by experts across many subjects.
Center for AI Safety, Scale AI · 4 models · 2 tasks
Building features in real software projects
113 original tasks across 91 real code projects in five languages, checked with hand-written tests.
Datacurve · 8 models · 1 task
Getting jobs done on a computer's command line
Practical tasks completed by typing commands, like setting up software or processing files.
Terminal-Bench team · 2 models · 1 task
Real paid freelance projects
Projects originally done by freelancers for pay, scored on whether the AI's result would be acceptable to the client.
Scale AI, Center for AI Safety · 2 models · 1 task
Research-level maths
New, unpublished maths problems that take specialists hours or days.
Epoch AI · 11 models · 1 task
Would a real maintainer accept its code?
Coding tasks in 36 major open-source projects, judged on whether the code is good enough to merge, not just whether it runs.
Cognition · 9 models · 1 task
People's votes
People ask a question, see two anonymous answers and pick the better one. It shows which answers people prefer, not whether they were right.
People's votes on writing and language
People compared two anonymous answers to writing, editing and language questions.
LMArena · 14 models · 9 tasks
People's votes on business and finance questions
People compared two anonymous answers to business, management and finance questions.
LMArena · 14 models · 6 tasks
People's votes on expert questions
People compared two anonymous answers to questions needing expert knowledge.
LMArena · 14 models · 6 tasks
People's votes on following instructions
People compared two anonymous answers to requests with specific instructions.
LMArena · 14 models · 6 tasks
People's votes on creative writing
People compared two anonymous pieces of creative writing.
LMArena · 14 models · 4 tasks
People's votes on back-and-forth conversations
People compared two anonymous answers in conversations with several back-and-forth turns.
LMArena · 14 models · 3 tasks
People's votes on health questions
People compared two anonymous answers to medical and healthcare questions.
LMArena · 13 models · 3 tasks
People's votes on long, detailed requests
People compared two anonymous answers to long requests with lots of detail to take in.
LMArena · 14 models · 3 tasks
People's votes on science questions
People compared two anonymous answers to questions in the life, physical and social sciences.
LMArena · 14 models · 3 tasks
People's votes on maths questions
People compared two anonymous answers to maths questions.
LMArena · 13 models · 2 tasks
People's votes on questions in other languages
People compared two anonymous answers to questions asked in languages other than English.
LMArena · 14 models · 2 tasks
People's votes on coding questions
People compared two anonymous answers to their coding questions and picked the better one.
LMArena · 14 models · 1 task
People's votes on legal questions
People compared two anonymous answers to legal and government questions.
LMArena · 14 models · 1 task
People's votes overall
People compared two anonymous answers to all kinds of questions and picked the better one.
LMArena · 14 models · 1 task