Asher Cohen
Back to posts

Did OpenAI Really Show Us How Capable AI Is? Or How Good It Is at Marketing Itself?

AI can contribute to serious mathematical research, but that is not the same as proving it can autonomously solve problems humans could not.

Did OpenAI Really Show Us How Capable AI Is? Or How Good It Is at Marketing Itself?

There is something uncomfortable about the way OpenAI has been presenting some of its recent mathematical achievements.

The underlying technology is clearly becoming extraordinarily capable. I don't think there is much value in denying that. AI systems are solving problems that were previously out of reach for machines, generating useful mathematical ideas, and increasingly contributing to genuine research.

But there is a significant difference between saying:

"An AI system contributed to solving a difficult mathematical problem."

and saying:

"The AI solved a problem that humans couldn't solve."

That difference matters enormously.

The recent controversy around GPT-5 and Erdős problems is a good example of why.

The Impressive Claim

OpenAI presented results suggesting that GPT-5 had solved a number of previously open mathematical problems associated with Paul Erdős.

These aren't ordinary textbook exercises. Erdős left behind a huge collection of mathematical problems, many of which have remained unresolved for decades.

The headline was therefore spectacular: an AI model had apparently succeeded where mathematicians had failed.

It's exactly the kind of result that makes for compelling headlines — and compelling evidence for the argument that AI is rapidly approaching, or even surpassing, human intellectual capabilities.

There was just one problem.

The details mattered.

What Happens When Humans Look More Closely?

When the results were examined more carefully, it became apparent that the story was considerably more complicated.

Some of the supposed solutions were connected to results that already existed. In other cases, the model's output required substantial human intervention: correcting mistakes, filling gaps, interpreting the argument, or turning an incomplete line of reasoning into a valid mathematical proof.

That doesn't make the model useless.

Quite the opposite.

Producing an idea that allows a mathematician to solve a problem can be an extremely valuable contribution.

But it does change the claim.

If a model generates a promising but flawed proof, a mathematician identifies the flaw, fixes it, completes missing steps and verifies the final result, I don't think it is intellectually honest to simply say:

"The model solved the problem."

A more accurate description would be:

"The model generated ideas that helped mathematicians solve the problem."

Those are very different achievements.

The Human Editing Problem

This is the part of the story I find most important.

Consider a simplified process:

  1. AI generates possible approaches
  2. AI finds something interesting
  3. Human identifies errors
  4. Human fills missing steps
  5. Human modifies the argument
  6. Mathematicians verify the proof
  7. Correct mathematical result

At the end of that process, we have a correct solution.

But who solved the problem?

The answer isn't necessarily "the AI."

The AI may have made a genuinely important contribution. It may even have provided the key insight.

But the final achievement belongs to a human + AI system, not necessarily to the underlying model operating autonomously.

That distinction becomes especially important when these results are used as evidence for claims about general AI capabilities.

Capability Versus Autonomy

This is where I think the current AI narrative often becomes confused.

A model can be extremely capable without being autonomous.

It can produce extraordinary ideas while still being unreliable.

It can solve a problem while failing to recognize when its own solution is wrong.

It can generate a proof that looks convincing but contains a subtle logical error.

And it can sometimes produce a brilliant answer after thousands of attempts while failing repeatedly on a superficially similar problem.

Those are not necessarily contradictions.

They are characteristics of a system that is powerful but still fundamentally probabilistic and unreliable.

The distinction is between:

"The model can produce a correct solution."

and:

"The model can reliably conduct mathematical research."

The second claim is much stronger.

The Benchmark I'd Actually Like to See

This is why I find benchmarks such as First Proof more interesting.

Rather than asking whether an AI system can produce an impressive result under favourable conditions, the question becomes something closer to:

Can AI solve genuinely new mathematical research problems?

In the First Proof evaluation, ten new mathematical problems were given to several AI systems. The strongest system reportedly solved six of them, while humans still performed better overall.

That's an impressive result.

But it is also a much more useful one.

It tells us something about the actual boundary of the technology.

AI is clearly capable of contributing to serious mathematical research.

It is not yet evidence that AI has become an autonomous mathematician.

Those two conclusions can coexist.

The Statistical Trick Hidden Inside Spectacular Results

There is another issue that gets very little attention in AI marketing: selection bias.

Suppose a model generates 10,000 possible approaches to a problem.

Imagine that 9,997 are useless and three contain genuinely valuable ideas.

One of those three eventually leads to a correct solution.

The statement:

"AI found a solution to an unsolved mathematical problem."

may technically be true.

But it tells us almost nothing about the model's actual reliability.

We don't know:

  • How many attempts were required
  • How much compute was used
  • How many incorrect solutions were discarded
  • How much human intervention was required
  • How often the model failed
  • Whether it could recognize its own failures
  • Whether the result could be reproduced
  • Whether the same methodology works on other problems

These are not minor implementation details.

They are the difference between demonstrating a possibility and demonstrating a capability.

"Something worked once" is not the same as "we can reliably delegate this task to the system."

This Doesn't Mean the AI Isn't Genuinely Impressive

I don't want to overcorrect here.

There is a temptation to look at cases like this and conclude that the whole thing is marketing hype.

I don't think that's right either.

The underlying progress is real.

Modern AI systems can produce mathematical ideas that are genuinely useful. They can explore enormous numbers of possibilities, connect concepts in unexpected ways, assist with proofs and sometimes contribute to results that humans could not easily obtain on their own.

That's remarkable.

The problem is not that the technology is incapable.

The problem is that the maximum demonstrated capability is often presented as if it were the normal, autonomous capability of the model.

Those aren't the same thing.

The System Is Not the Model

This distinction is particularly important.

When we say "GPT solved the problem," what exactly are we talking about?

  • The neural network?
  • The model plus inference-time search?
  • The model plus external tools?
  • The model plus thousands of generated attempts?
  • The model plus a human mathematician?
  • The model plus a human who knows which outputs are promising?
  • The model plus a formal verification system?

All of these can produce very different results.

Calling all of them simply "the model" makes comparisons almost meaningless.

It would be similar to evaluating a software engineer by giving them an IDE, compiler, debugger, Stack Overflow, a team of reviewers and an experienced architect — and then attributing the entire result to the programmer alone.

The system matters.

The workflow matters.

The humans matter.

The amount of search matters.

And the evaluation methodology matters.

The Real Breakthrough Would Look Different

For me, the convincing experiment would be much simpler.

Take ten genuinely new mathematical problems.

Give them to the model.

Specify the computational budget.

Allow no human intervention.

Require the model to produce complete proofs.

Have those proofs independently verified, ideally formally.

Then repeat the experiment across many different problems.

If an AI system consistently solved, say, eight or nine out of ten difficult research problems under those conditions, that would be extraordinary evidence.

At that point, I would have no problem saying that we were witnessing something much closer to autonomous mathematical research.

We're not there yet.

And pretending that we are doesn't help.

The Irony

The irony is that we don't actually need to exaggerate these systems.

The genuine achievements are already impressive enough.

AI doesn't need to have autonomously solved every Erdős problem to be revolutionary.

A system that can generate an insight that a mathematician can turn into a new theorem is already incredibly useful.

A system that can search mathematical spaces humans cannot feasibly explore is already valuable.

A system that can accelerate research by an order of magnitude would fundamentally change science, even if a human remains responsible for verification.

That is already a remarkable story.

We don't need to turn it into:

"AI has replaced mathematicians."

And this is where I think OpenAI's communication deserves criticism.

The issue isn't necessarily that the underlying results were fake.

The issue is that the framing can blur the boundary between what the AI generated, what the complete human-AI system achieved, and what humans ultimately verified and corrected.

That creates a distorted perception of AI capability.

We Should Be Careful with the Word "Solved"

This may seem like semantics.

It isn't.

In mathematics, "solved" has a very specific meaning.

A proof either establishes the claim or it doesn't.

And when we're evaluating an AI system, there is another question:

Did the AI actually discover the proof, or did humans turn the AI's output into a proof?

Those questions need separate answers.

Otherwise we end up measuring the capabilities of an entire human-AI workflow while claiming to measure the intelligence of the model.

That is where marketing starts to replace measurement.

The Conclusion I'd Take from All This

My takeaway isn't that AI is less capable than people claim.

It's that we need better ways of measuring what "capable" actually means.

The current generation of models is astonishingly good at producing useful intellectual work.

But impressive outputs are not enough.

We need to know the failure rate, the amount of search, the human involvement, the reproducibility and the degree of autonomy.

Until then, I would be very careful with claims that an AI system has "solved" a problem that humans couldn't.

The more accurate description is often more interesting anyway:

AI is becoming an extraordinarily powerful research tool, but we haven't yet demonstrated that it can reliably conduct research on its own.

That's a much less sensational headline.

It is also, at least for now, a much more defensible one.

#ai #research #mathematics #benchmarking