Can't play sound? The illustrated version is right below.
The full narration of the video.
You might think if a computer says it solved a hard math problem, then it's solved. It's not that simple. On October sixth OpenAI put a huge stack of math papers online. 722 papers, all written by an AI that hasn't been released.
They gave this AI about four thousand problems all ones mathematicians had never solved. On average, each result took it about three hours of thinking. Some of these questions are more than a hundred years old. Take this one. A needle has to turn all the way around on a patch of ground. How small can that patch be?
Someone asked this in nineteen seventeen. On a flat surface, the answer is a surprise the patch can have almost no area at all. But in three dimensions, a new question came up however small the space gets, can it ever be as flat as a sheet of paper? Human mathematicians only proved it last year it can't. One paper in this stack says it has proved the same thing with one more dimension.
So, is it solved? Not so fast. Think of an exam. Handing it in is quick, grading is the hard part. Now more than seven hundred answer sheets have landed on the desk at once. Luckily, there's a helper. A computer program built to check math proofs. Every step must be spelled out, skip one, and it won't pass.
It's called Lean. Of the 700 or so papers, 300 have their main result through its check. But there's something it can't catch it checks every step, not whether the question was copied right. If the question is a little off, every step can be right and still answer a different question. Whether it's new, whether it matters, people still have to judge.
The next day, October seventh, OpenAI posted a correction. In one paper, a plus or minus sign was flipped. That one mistake pulled three papers back. Fourteen others had their proofs repaired. Mathematicians say it will take months to tell which of these are truly new.
Some say this is good for math. Others say show us the receipts. How the AI did it, and the exact questions it got, all out in the open. So when a computer says it solved it, does that count? Right now the answer is the paper is handed in, and grading has just begun. Math used to be hard because proofs were hard to write. Soon the hard part may be finding enough people to read them.
Which hard problem would you want an AI to take on? Tell us in the comments.
You might think: if a computer says it solved a hard math problem, then it's solved.
It's not that simple. On October 6 (US time), OpenAI put a huge stack of math papers online: 722 of them, all written by an AI it hasn't released.
They gave this AI about 4,000 problems mathematicians had never solved. On average, each result took it about three hours of thinking.
A needle has to turn all the way around on a patch of ground. How small can the patch be? Someone asked this in 1917. On a flat surface the answer is a surprise: it can have almost no area at all.
In three dimensions, a new question came up: however small the space gets, can it ever be as flat as a sheet of paper? Human mathematicians only proved the answer in 2025: no.
One paper in this stack says it proved the same thing with one more dimension.
Think of an exam: handing it in takes minutes, grading can take all day. Now more than seven hundred answer sheets have landed on the desk at once.
There's a computer program built to check math proofs. It's called Lean. By October 7, 300 of the 719 papers had their main result through its check, about four in ten.
It only checks each step. If the question is copied a little wrong, every step can be right and still answer a different question.
Whether a result is new, and whether it matters, people still have to judge.
On October 7, OpenAI posted its own correction: in one paper, a plus or minus sign was flipped. That one mistake pulled three papers back, and fourteen others had their proofs repaired.
Mathematicians say it will take months to tell which of these are truly new. Some think putting them out in public is good for math. Others say: show us the receipts. How the AI did it, and the exact questions it got, all out in the open.
When an AI keeps failing in the same place, the fix is often not the AI: let a second AI review old mistakes and suggest edits, and a human approves.
The AI helper you hired got a new version, and nine of its habits changed. A symptom chart: what feels off, and the one line to say for each.
A new AI that never talks: it fills in bubbles on a printed answer sheet with a confidence note. Why it is 40-200x faster, and why it can't make things up.
Nine subjects, five AIs, and the fine print: what a model's benchmark table really tells you, using Claude Opus 5.5 as the example.