Five AIs took the same exam in nine subjects. The newcomer came first in seven of the nine.
The company that built it published this table itself, like a student posting their own report card. The scores come from real tests, but which subjects to show and how to lay them out was its own choice.
2
Who took the exam
Opus 5.5 the newcomer, today's star made by Anthropic
Fable 5.1 the priciest, strongest one from the same maker Anthropic
Opus 5 its previous generation, to show the progress Anthropic
GPT-6 Astra the rival's strongest made by OpenAI
GPT-5.6 Sol another one from the rival OpenAI
A "—" in the table means that AI skipped that exam. It did not score zero.
3
How this exam differs from a normal exam
TESTING AI BEFORE
One question, and it writes one answer
Like a written test: right or wrong, judged on the words
TESTING AI NOW
Hand it a computer and a job, and let it get the job done itself
Like a practical exam: it only scores if what it hands in actually works
Several subjects in the table say "agentic", which means "let it do the work itself", not just answer questions.
4
Nine subjects, one by one
The big number is Opus 5.5's score. On the two gray cards, it did not come first.
Working alone at the command line
66.4
Terminal-Bench 4.0
It only gets a window for typing commands to the computer, and has to finish the job one command at a time. Out of 100 jobs, it got 66 done.
First, ahead of second place by more than 8 points, the widest gap of any subject
Hard programming jobs
54.4
FrontierCode
Real tasks that even professional programmers have to chew on for a while.
First, but the rival scored 53.3, just one point behind, so basically a tie
Working inside a coding app
57.8
CursorBench
A coding app programmers use every day, tested by its maker on real work.
First
Real office work
1846
GDPval-AA
The kind of documents lawyers, accountants and analysts hand in. Two AIs each make one, and a person picks the better one. This is not a 100-point scale but a chess-style rating: the more you win, the higher it goes.
First, 138 points above the previous generation
Chaining office apps into a workflow
40.0
AutomationBench
For example: "get an email → fill in a spreadsheet → notify a coworker", a whole chain of steps.
Lost: the rival scored 41.4. This exam was run by another company, and if a safeguard blocks it, that counts as a failure (see part 5)
The hardest questions in every field
67.7
Humanity's Last Exam
Questions written by experts around the world specifically to stump AI. Looking things up and using a calculator are allowed.
First
Doing research on its own
58.7
Terminal-Bench-Science
It designs the experiments, runs them and reports the results by itself.
Lost: the rival scored 64.6. But it doubled its previous generation's 29.0
Using a computer like a person
81.8
OSWorld
Looking at the screen, clicking the mouse, typing and opening apps to finish a task.
First
Reading charts
89.0
Chartography
Show it a bar chart or line chart and ask what the numbers are and which one grew fastest.
First, but second place scored 88.4, less than a point behind, so a tie
5
The fine print under the table matters more than the scores
Tested at "full effort"By default it runs at medium effort. In everyday use, you usually won't reach the scores in the table.
Tested with its safety lock onOn sensitive topics like cyber attack and defense or biology, the lock hands the question to an older version. So a few scores were pulled down by its own lock.
Every score wobbles by 2 to 5 pointsRetake the same exam and the score moves up or down a few points. A one or two point lead is not a real win. Ten points is.
6
What a report card can't tell you
A good test score doesn't mean it does your job wellAll nine subjects were chosen by whoever set the exam. Writing copy, making spreadsheets, replying to customers: not one of those is in the table.
"So should I switch to it?"
The table says it improved most at jobs it does by itself: writing code and operating a computer. But the best exam is your own work: give old and new the same task and see which result you'd rather use.