← ELI5 中文
EXPLAINED SIMPLY

An AI handed in 700+ math papers. Solved?

Handing in is easy. Grading is hard.
For anyone who's curious
No. 129How AI WorksPart 15 of 15Fact-checked Oct 8, 2026
Video2:34 · English voice (AI) · English subtitles

Can't play sound? The illustrated version is right below.

Read the transcript

The full narration of the video.

You might think if a computer says it solved a hard math problem, then it's solved. It's not that simple. On October sixth OpenAI put a huge stack of math papers online. 722 papers, all written by an AI that hasn't been released.

They gave this AI about four thousand problems all ones mathematicians had never solved. On average, each result took it about three hours of thinking. Some of these questions are more than a hundred years old. Take this one. A needle has to turn all the way around on a patch of ground. How small can that patch be?

Someone asked this in nineteen seventeen. On a flat surface, the answer is a surprise the patch can have almost no area at all. But in three dimensions, a new question came up however small the space gets, can it ever be as flat as a sheet of paper? Human mathematicians only proved it last year it can't. One paper in this stack says it has proved the same thing with one more dimension.

So, is it solved? Not so fast. Think of an exam. Handing it in is quick, grading is the hard part. Now more than seven hundred answer sheets have landed on the desk at once. Luckily, there's a helper. A computer program built to check math proofs. Every step must be spelled out, skip one, and it won't pass.

It's called Lean. Of the 700 or so papers, 300 have their main result through its check. But there's something it can't catch it checks every step, not whether the question was copied right. If the question is a little off, every step can be right and still answer a different question. Whether it's new, whether it matters, people still have to judge.

The next day, October seventh, OpenAI posted a correction. In one paper, a plus or minus sign was flipped. That one mistake pulled three papers back. Fourteen others had their proofs repaired. Mathematicians say it will take months to tell which of these are truly new.

Some say this is good for math. Others say show us the receipts. How the AI did it, and the exact questions it got, all out in the open. So when a computer says it solved it, does that count? Right now the answer is the paper is handed in, and grading has just begun. Math used to be hard because proofs were hard to write. Soon the hard part may be finding enough people to read them.

Which hard problem would you want an AI to take on? Tell us in the comments.

You might think: if a computer says it solved a hard math problem, then it's solved.

It's not that simple. On October 6 (US time), OpenAI put a huge stack of math papers online: 722 of them, all written by an AI it hasn't released.

They gave this AI about 4,000 problems mathematicians had never solved. On average, each result took it about three hours of thinking.

PICK ONE

Turn a needle around

A needle has to turn all the way around on a patch of ground. How small can the patch be? Someone asked this in 1917. On a flat surface the answer is a surprise: it can have almost no area at all.

In three dimensions, a new question came up: however small the space gets, can it ever be as flat as a sheet of paper? Human mathematicians only proved the answer in 2025: no.

One paper in this stack says it proved the same thing with one more dimension.

NOT SO FAST

Handing in is quick.
Grading takes work.

Think of an exam: handing it in takes minutes, grading can take all day. Now more than seven hundred answer sheets have landed on the desk at once.

A HELPER

A grader that
only trusts steps

Every step spelled outIf each step follows, it passes.
Skip a stepMiss one, and it fails.

There's a computer program built to check math proofs. It's called Lean. By October 7, 300 of the 719 papers had their main result through its check, about four in ten.

300 passed719 papers
WHAT IT CAN'T CATCH

It never checks
the question

Question

It only checks each step. If the question is copied a little wrong, every step can be right and still answer a different question.

Whether a result is new, and whether it matters, people still have to judge.

THE NEXT DAY

One flipped sign,
three papers pulled

±

On October 7, OpenAI posted its own correction: in one paper, a plus or minus sign was flipped. That one mistake pulled three papers back, and fourteen others had their proofs repaired.

WHAT MATHEMATICIANS SAY

It will take
months to tell

Mathematicians say it will take months to tell which of these are truly new. Some think putting them out in public is good for math. Others say: show us the receipts. How the AI did it, and the exact questions it got, all out in the open.

A computer says
it solved it. Does it count?
Handed in, still being graded.
Math used to be hard because proofs were hard to write. Soon the hard part may be finding enough people to read them. Which hard problem would you want an AI to take on?
UP NEXT

Let It Rewrite Its Own Manual

When an AI keeps failing in the same place, the fix is often not the AI: let a second AI review old mistakes and suggest edits, and a human approves.

No. 8 · 4 minKeep going →
← Back to ELI5