OpenAI Math Release: What’s Real in the 722 AI Manuscripts

11 min read

Online Tech Tips is reader-supported. We may earn a commission when you buy through links on our site. Learn more.

On October 6, 2026, OpenAI published a post on its official blog called “Sharing AI progress in mathematics.” It initially released 722 math manuscripts, grouped into 372 “result families,” and said an unreleased internal AI model produced them. A result family can bundle related results, alternative proofs, and dependent manuscripts, so 719 manuscripts is not the same as 719 separate breakthroughs. Three were withdrawn the next day over a sign error, so the catalog now lists 719 manuscripts (TechCrunch). It’s a big claim. Some of the work is checkable, some has already been pulled back, and much of it is still waiting for a human to read it.

So does this mean AI now beats human mathematicians or rival AI systems at math? No. It isn’t a new ChatGPT feature you can switch on today, either. What it does suggest, in our view, is that an AI lab can now produce math research faster than the math world can check it. (All sources were retrieved on October 10, 2026. Claims about the manuscripts come from OpenAI unless we say otherwise.)

OpenAI blog post Sharing AI progress in mathematics with its October 6, 2026 date

What Happened

According to OpenAI’s announcement, the company published a large batch of math write-ups made by what it calls an internal frontier model. That model hasn’t been released to the public. The topics range widely:

  • Number theory, algebra, geometry, and topology
  • Probability and mathematical physics
  • Quantum algorithms and theoretical computer science

OpenAI says the model took on hundreds of open or unsolved math questions. Some results touch on very famous problems. These include the Riemann hypothesis, a special case of the Hodge conjecture, and the four-dimensional Kakeya conjecture.

Notice the wording there: “touch on.” That’s a much weaker claim than “solved.” More on that in a second.

A few other details stand out:

  • Unequal confidence levels: OpenAI labeled results as sitting at different stages of verification, per The Verge. It didn’t present everything as fully confirmed.
  • Only some proofs are formalized in Lean: Some proofs were written in Lean, a language that lets a computer check a formal proof step by step. TechCrunch reports that roughly 300 of the 719 top-line results (about 42%) were formalized in Lean, and that OpenAI shared only 10 summaries of the model’s reasoning (TechCrunch).
  • Lots of computing power: An average result took about three hours of computing time, a figure reported by Fortune’s report.
  • Code and files in public: The Verge covered the materials OpenAI posted on GitHub, so anyone can browse the manuscripts.

How We Got Here: A Quick Timeline

This didn’t come out of nowhere. The October release follows a rocky September.

  • September 2026: The open-problem claims. OpenAI claimed its model had resolved more than 100 longstanding open problems, a self-reported count that had not been independently checked.
  • September 8, 2026: The Navier–Stokes claim. OpenAI announced that it had solved a version of the Navier–Stokes problem (the case with a smooth external force), releasing a 166-page paper and a Lean formalization (Wikipedia’s summary of the dispute). These equations describe how fluids like water and air move, and the problem is one of the Clay Millennium Prize Problems. The claim drew pushback over rushed announcements and credit for prior human work, and more than 8,000 researchers endorsed those concerns, per Retraction Watch. The Clay Mathematics Institute says its verification process is “deliberately unhurried” and lists the problem as “active,” not solved, and a later paper by UK mathematicians flagged places where the written proof and its Lean code diverge (TechCrunch).
  • Before October: Outside reviewers join. After that criticism, OpenAI said it worked with an independent math advisory group hosted by the Institute for Advanced Study.
  • October 6, 2026: The big drop. 722 manuscripts, 372 result families, all at once. ThePrint described it as OpenAI dropping “the mother lode on mathematicians.”
  • October 7, 2026: Three manuscripts withdrawn. A sign error broke the argument in one manuscript and constructions used by two papers that depended on it, so all three were withdrawn. OpenAI also revised 14 other manuscripts with proof repairs and corrected statements, per Retraction Watch. That leaves 719 in the catalog.

Who This Affects

Let’s clear up the biggest confusion first.

Is this a new ChatGPT feature I can use? No. The model behind these manuscripts is internal and unreleased. If you use ChatGPT on the web, on your phone, or through a Plus or Pro plan, nothing changed for you. There’s no new button and no new “math mode.” You can read the blog post and browse the files for free, with no account and no app.

The people this does affect right now are:

  • Professional mathematicians, who are being asked to check a huge pile of new work.
  • Grad students and early-career researchers, who worry about what happens to their research areas (more on that in The Reaction).
  • Students and teachers, who want to know if AI can now “do math” reliably. The honest answer: it can produce impressive leads, but leads aren’t proofs until someone checks them.
  • Anyone following the AI race between OpenAI, Google, and Anthropic.

Why “AI Solved 372 Problems” Is Messier Than It Sounds

Did OpenAI really solve the Riemann hypothesis?

No, and OpenAI hasn’t claimed it did. The company says some results touch on the Riemann hypothesis and other famous problems. That usually means work on a related piece, a special case, or a side question. Think of it like fixing a leaky faucet in a famous building. It’s real work, but you didn’t renovate the whole building.

Why are some solutions being withdrawn?

Because math is unforgiving. One wrong step anywhere breaks the whole proof. A single sign error knocked out three manuscripts, because two of them built on the first. That’s the system working the way it should. Still, it shows that “published” and “correct” are two very different things.

TechCrunch’s coverage made a similar point. Its report says the solutions aren’t meeting the field’s standards yet.

Why does “formalized in Lean” matter?

Lean is like a super-strict spell checker for logic. You write a proof in Lean’s code-like language, and the computer checks every step. If Lean accepts it, the formal argument follows from the formal premises. That’s not the same as “verified” in the everyday sense, though.

So why wasn’t everything written in Lean? Converting a proof into Lean is slow, hard work. Lots of modern math also doesn’t have the building blocks it needs in Lean yet. Only about 42% of the top-line results were formalized, according to TechCrunch.

There’s a sneakier catch, too. Lean checks the statement as it was typed into Lean. If that statement doesn’t quite match the English write-up, the computer can approve something narrower than the headline. That gap is a big reason mathematicians still want human eyes on all of this. This isn’t hypothetical: the UK paper TechCrunch described found exactly this kind of mismatch in the Navier–Stokes work. A green checkmark from Lean is great news. It just has to be a checkmark on the right question, and it says nothing about whether a result is new or important.

So what’s actually verified?

Here’s where things stand, based on what OpenAI and independent outlets have reported:

ClaimStatusSource
OpenAI published 722 manuscripts in 372 result familiesConfirmed: 722 initially, 719 after three withdrawalsOpenAI, TechCrunch
An unreleased internal model produced themOpenAI’s own claim; the model isn’t publicOpenAI
All results are correctNot verified. OpenAI itself labels them as being at different stages of verificationThe Verge
All results are formalized and checked in LeanNo. Roughly 300 of 719 top-line results (about 42%) were formalized in LeanTechCrunch
All results are peer-reviewedNo. This is a self-published research releaseTechCrunch
Some results were withdrawnYes: three on October 7 over a sign error, plus 14 revisedRetraction Watch
Navier–Stokes is solvedNot officially. OpenAI claims a solution with a Lean formalization; Clay lists the problem as “active” and verification is ongoingWikipedia

How It Stacks Up

Does this mean OpenAI’s AI is smarter at math than Google’s or Anthropic’s? We can’t say that, and neither can anyone else yet.

Here’s why. OpenAI tested its model on open research problems it picked. None of the coverage we reviewed reports a shared, published benchmark against the latest Claude or Gemini models, or any independent head-to-head comparison. The model is also private, so nobody else can test it.

It’s a bit like a runner announcing a great time on a course they designed, with nobody else on the track. The time might be great! Without other runners, though, nobody can call it a win.

OpenAI’s October 2026 releaseGoogle DeepMind math AI (e.g., AlphaProof-type systems)Anthropic’s Claude modelsTraditional peer-reviewed math
What’s on the table722 manuscripts, 372 result families (self-reported)Its own formal and informal math researchGeneral-purpose AI modelsPapers checked by independent experts
Head-to-head benchmark vs. the othersNone publishedNone matched to this releaseNone matched to this releaseNot applicable
Independent verificationPartial: about 42% of top-line results formalized in Lean, outside reviewers, three withdrawalsNot compared in current sourcesNot compared in current sourcesRequired before acceptance
Can you use it?No, model is unreleasedNot compared hereClaude is publicly availableNot applicable

Comparison based on the research sources cited above, retrieved October 10, 2026. “None” means no verified comparison appeared in those sources.

Bottom line: anyone telling you OpenAI is now “ahead” of Claude or Gemini at math is guessing.

Retraction Watch article headline about OpenAI withdrawing three preprints after releasing 722 math manuscripts

The Reaction

Mathematicians are split, but the split is mostly about process rather than whether the math is interesting. Supporters see real results worth checking. Critics argue that dumping hundreds of unreviewed manuscripts at once shifts the cost of verification onto unpaid human experts. Most of the on-record comments land somewhere between those two positions.

  • Enthusiasm with a warning: Dan Litt of the University of Toronto told Fortune, “My view is that this is great for mathematics,” while urging support for human mathematical expertise so funding doesn’t dry up.
  • “Math 1.0” is over: Terence Tao of UCLA argued that “problems are being solved autonomously by AI prompters who have no interest in the broader field,” as quoted by TechCrunch and Fortune.
  • The checking burden: Harvard’s Melanie Wood said that when models solve problems, “there is not human understanding of them at the point of release, and now the work begins” (TechCrunch).
  • Release order: Cornell’s Alex Townsend told Retraction Watch that OpenAI “should have announced the manuscripts that were lean verified first.” MIT’s Andrew Sutherland called the quick withdrawals “the responsible thing to do,” but said many mathematicians remain “unhappy with how OpenAI has behaved up to this point.”
  • Harsher critics: NYU’s Tristan Buckmaster told the New York Times that OpenAI likely hadn’t done its “due diligence,” per Fortune, and a group called the Association for Human Mathematics described the release as “not a demonstration of scholarship, but a demonstration of power.”

The IAS-hosted advisory group OpenAI worked with has also published recommendations for responsible AI-proof publication. TechCrunch reports the release falls short of several of them, including the first one: stop testing proprietary models on open problems.

Our read: the backlash makes sense. OpenAI handed the hard part, checking, to everyone else.

Our Take

This is a real and impressive research release. Still, it has no scoreboard, and there’s nothing here you can use.

Here’s who should care, and why:

  • Everyday ChatGPT users: Nothing changes for you today. Don’t upgrade your plan expecting this model. It isn’t in ChatGPT.
  • Students: Treat AI math output like a smart classmate’s notes. It’s useful for ideas, but always double-check the steps. Even OpenAI’s top internal model produced results that got pulled within days.
  • People following the AI race: File this under “promising, unproven.” Without a shared benchmark, it says nothing reliable about OpenAI versus Google or Anthropic.
  • Mathematicians and educators: This is the group that should care most. The flood of model-generated work is real, and the checking process wasn’t built for it.

We’ll give OpenAI credit for bringing in outside reviewers after the Navier–Stokes dispute, for formalizing part of the work, and for withdrawing flawed papers quickly. But releasing 722 manuscripts at “different stages of verification” puts speed ahead of certainty. In math, certainty is the whole point.

What To Do Next

Want to judge it for yourself? Read OpenAI’s announcement first, then The Verge’s coverage right after. Putting the company’s claim next to independent reporting is the fastest way to spot the gap between the headline and what’s verified.

Thinking about a paid AI plan for homework or work math? ChatGPT Plus, Claude Pro, and Google’s Gemini plans are all options. Pick based on what those public models do today, not on this unreleased one.

Wrapping Up

OpenAI dropped a mountain of model-generated math on October 6, 2026, and it’s genuinely remarkable. But it’s mixed-confidence, only partly formalized in Lean, already trimmed to 719 after withdrawals, and not peer-reviewed. You can’t use any of it in ChatGPT.

Our honest take: the biggest lesson here is that AI can now create math faster than humans can check it. That’s a problem the whole field will have to solve next (hopefully without another Navier–Stokes-sized headache!).