Because they apparently lack the internal compentencies to do so (or, given that they now have a Fields medalist on payroll, it's probably more a case of not feeling any need whatsoever to spend the effort it takes; maths isn't exactly profitable if you can't mine it for PR). Ignoring all the ethical violations they made in the process, the fact that they only ever released slop on the Navier–Stokes problem is part of the debacle.
For example, had they decided to put their paper on arXiv, they would be facing a 1 year ban, followed by not being able to put anything on there without prior peer review.
More: after OAI've mangled whatever value is left of the results (once again), HN could start braying for the training data/weights. Or even the prompts
Given that he apparently feels exploited after an appearance in some of their earlier marketing material, I imagine he's not overly keen on working more for their PR department, or this extension thereof.
> During this event, OpenAI requested an interview concerning my vision of the future of AI and mathematics. I accepted, and spoke with them for perhaps an hour. I had done similar interviews in various venues, and I assumed that, as with these other cases, they would eventually post the entire interview online, which talked about both the possibilities and risks of AI much as I have done in these other interviews. As it turned out, they only used a few snippets of that interview for that infamous advertisement instead.
Lest someone gets the wrong impression, I think that any mention of the a priori strange-looking resolution to the Jacobian conjecture should be accompanied with a read of pages 160–161 of https://arxiv.org/abs/2609.05746. The tl;dr being that the very same construction featuring in the resolution appeared earlier in a draft paper that was accidentally made publically available for a travel award application. The same story has accompanied several other of the major announcements made by now. We haven't have a move 37 for maths yet, despite what OpenAI's marketing department might want you to believe.
Similarly, if all you ever read were OpenAI blog posts, you would get a very wrong impression of the usefulness of large language models of today in maths. For a working researcher, it's not a magic wand that you point at any given proposition and it tells you whether that proposition is true or not. It does appear to help if, while pointing your wand and utter the magical incantation “do it up bro”, you also make it convert $15 million into heat, but for most people, this kind of inverted Midas touch isn't quite accessible yet.
Instead, the reality seems to be closer to this, projecting a fair bit: a given mathematician will have a collection of propositions that they care about, and that they'll use as their own internal benchmark as new models come out. Very rarely will anything come out of it, but sometimes, in particular if you make sure to provide the wand with all relevant context, papers that could be relevant, proof strategies and lemma structures that you suspect are useful, something (which may or may not be plagiarism) will pop out, and that's really nifty. Moreover, it is not unimportant what the proposition and the relevant proof is like. And what does come out tends to be quite bizarre; proofs that use terminology that doesn't exist, seem overly pretentious, based on nonsense analogies where it's surprising that it even works at all, and the only comfort is that you can join it with an equally unreadable Lean blob. And where you would be _crazy_ to just publish those artifacts and think that you have contributed much of anything to maths.
But sometimes it works. It's still very unclear what kind of maths the models are good at, but it seems to certainly be an advantage if what you're looking for is a counterexample hidden in a pile of otherwise similar-looking non-counterexamples, if your proof is one that requires considering 36 different cases, each of which are so tedious that no researcher would have the patience to go through them by hand, or if the proof is an amalgamation of several existing structures, some of which are only documented in Georgian.
The gold rush, more than anything else, seems to be populating the convex hull of existing maths.
This can all change. The $15 million wand requirement today will be less tomorrow. Whether we ever get a move 37 is less clear, or whether we will eventually reach stagnation as all low-hanging fruit is picked, and the convex hull is populated; call this cope if you like. But maybe we do get move 37s all over the place, and it's fine that people think about what that future will look like.
Until then, and while we're still picking friut, let us rather have a think about what we can do to fix the incentive mismatch, to ensure that we increase the prestige of digestion over being the first to convince the LLM to do it up. Since that's the one thing everyone seems to agree, chances are it'll probably converge to something that doesn't have to be written in commandment form, but out of the guest posts hosted by Tao so far, the one by Antieau has some useful suggestions for standards (that aren't entirely unlike those from Leiden): https://terrytao.wordpress.com/2026/09/15/fast-math-slow-mat...
> It does appear to help if, while pointing your wand and utter the magical incantation “do it up bro”, you also make it convert $15 million into heat, but for most people, this kind of inverted Midas touch isn't quite accessible yet.
That's like where chess was when Deep Blue was built by IBM. Productivity improved. There was someone complaining on here recently that the seat-back entertainment system on some airline had a chess program set to "trounce all humans".
> but it is undoubtedly and objectively accelerating research.
Part of the point of the letter is that it is quite possible to act in a way that is a net negative to research. The most obvious case is when the companies violate ethical standards in research.
The subtler case, the one for maths in particular, is what happens when you fail to follow well-established patterns for making maths research productive. Tao himself spelled out how that can look in https://mathstodon.xyz/@tao/117207856734787448 (which notably came before any of the news on Navier–Stokes).
> The declaration literally opens by writing that AI companies (or really anybody) saying, “Hey, let’s see if this powerful reasoning engine can solve an open problem in mathematics” is “detrimental” to the “science” of mathematics. Full stop.
So let's just quickly agree that the actual quote is “However, the push by AI companies to solve mathematical problems as a benchmark is detrimental to the science of mathematics, and to the mathematical community.” And that in this, "as a benchmark" is load-bearing.
We wouldn't see nearly the same amount of contempt from researchers had OpenAI picked a research-friendly approach.
What they did: Hear a rumour about the problem being solved by other researchers, then rush to scoop them (unethical), then, when they actually go talk to them, they try to oust an author (also unethical), and when they finally decide to share their own work, do so in the least useful way possible.
What they could have done: Upon hearing the rumours, connect with the researcher in question and propose that they join efforts instead; set up a joint project to test if the machines are useful in any way, and if that's not appreciated, back down again. And instead of dumping only an undigested paper* and a Lean proof, do the digestion prior to publishing anything (as Buckmaster was in the process of doing). If their own lack of competences was keeping them from digesting it, then again, reach out to the researchers to understand if anyone would be willing to do so.
In the second of those two worlds, we wouldn't be seeing nearly the amount of outrage that we are seeing right now.
*: Here, “digestion” is the process of turning an AI slop paper into something humans can read. LLMs can indeed sometimes (if much more rarely than marketing material from the large LLM companies will suggest) produce correct proofs, but they are often written in bizarre ways – they'll use lingo that doesn't exist, seem overly pretentious, dwell on extremely easy steps while glossing over the hard ones. Currently, a real researcher will take that output and transform it into something that others can understand, use, and build upon. This is not so different from what happens when using it to write software, although as someone who does both, I will say that the amount of digestion needed for proofs tends to be orders of magnitudes larger than for code. This meme is quite accurate: https://mathstodon.xyz/@tao/117068266071803252
There are several cases of this already. Bubeck himself had to retract earlier claims of novelty, and more recently, the provenance of the non-sofic group result was brought into question. Most recently, it turn out that the construction used for Anthropic's counterexample to the Jacobian Conjecture had appeared in an unpublished but publically available draft: https://news.ycombinator.com/item?id=49657499
This has a few practical implications: First of all, if you are in the target group of the marketing material, be wary. While these things can do non-trivial stuff, the amount of magic is being grossly over-stated. But also, when several of the big results have indeed been reappropriating the work of others; when the companies fail to provide proper attribution (the NS case in particular is laughable) and present the results as the models' own work, that's plagiarism.
And just to spell it out, since it looks like HackerNews is flooded by people who are new to science these days: even if a result doesn't come with a price, scholarly peer review is the norm across all of science: https://en.wikipedia.org/wiki/Scholarly_peer_review
For example, had they decided to put their paper on arXiv, they would be facing a 1 year ban, followed by not being able to put anything on there without prior peer review.
reply