OpenAI says it cracked one of math’s grand challenges. But there are troubling questions about how they did it—and what it means for us all
OpenAI says it cracked one of math’s grand challenges. But there are troubling questions about how they did it—and what it means for us all
Hello and welcome to Eye on AI. In this edition:
OpenAI claims it made a mathematical breakthrough. But some mathematicians raise questions about cheating—and intimidation.
Google DeepMind uses AI to predict the impact of genetic mutations.
OpenAI agents swarmed a German wiki—and OpenAI stayed quiet about it.
Mistral valued at $24.4 billion in new fund raise.
Google DeepMind examines why AI agents cheat.
Average Americans are pessimistic about AI’s impacts.
Apologies in advance for a long essay today. But there are several important points to be made and the background is, well, complicated.Over the weekend, rumors swirled that Anthropic was on the cusp of announcing that one of its AI models had cracked one of the Millennium Prize Problems. These are seven complex mathematical challenges that the Clay Mathematics Institute, founded by American mutual fund magnate Landon Clay, selected in the year 2000, offering a $1 million prize for the first correct solution to each problem.The specific problem that Anthropic had cracked, the rumors said, was something called the Navier-Stokes equations. These come from the field of physics, where they explain certain properties in fluid dynamics, and are useful for everything from weather forecasting to aircraft design. For everyday, empirical purposes, the equations work well, but mathematicians have never been able to prove whether the equations hold for all fluid interactions across all time sequences. Are there are special circumstances under which the equations break down, resulting in what is known as a “singularity”: a point at which one or more fluid properties, such as pressure or velocity, “blow up”—i.e. race off to infinity? Proving that such singularities exist or that the equations hold for all conditions is what the challenge is all about.Now, as I write this on Tuesday, we’ve learned a bit more about what happened—and the story turns out to be more complicated, controversial, and acrimonious than simply being the case that one of Anthropic’s AI models has solved Navier-Stokes, which it turned out it did not. Instead, OpenAI today announced that a multi-agent system, powered and coordinated by an unreleased internal model, and which at one point had 10,000 different sub-agents working different parts and variations of the problem, has solved Navier-Stokes. OpenAI’s AI proved that, in fact, there are conditions under which the equations will “blow up.” Yet, how exactly OpenAI came to solve Navier-Stokes is, it turns out, a matter of great controversy.
Mathematician questions how OpenAI hit upon its approach
In short: Tristan Buckmaster, a well-regarded mathematician at New York University’s Courant Institute, also released a statement prior to OpenAI’s saying that he and Levent Alpöge, a mathematician who works for Anthropic, used several different AI models from both Anthropic and OpenAI to discover an almost identical solution to one portion of the Navier-Stokes Millennium Prize problem—although they did not have a proof for the entire problem.
Buckmaster says that he and Alpöge took a concept for tackling the Navier-Stokes problem that had been pioneered by two other mathematicians, Diego Cordoba and Luis Martinez-Zoroa, and then used Anthropic’s Claude and OpenAI’s Codex powered by the GPT-5.6 Sol model, to push Cordoba and Martinez-Zoroa’s lines of attack through to completion. (Buckmaster said they also used OpenAI’s new Astra model to help them audit and write up their results but not for the actual mathematical reasoning and calculations.) Buckmaster says that he and Alpöge worked for most of a year, making only slow progress, but that with help from several AI models, they made rapid progress from mid-August onwards. He calls this “a Deep Blue-Kasparov” moment for mathematics (referring to the 1997 contest in which a computer chess program first defeated a human grandmaster) and says “the significance of this with respect to the way we train students, assign credit, referee, and decide what is worth one human life’s attention cannot be understated.” (We’ll get back to this theme later.)
Then, however, Buckmaster made a series of explosive revelations. He said OpenAI had desperately asked for a phone call with him, starting on September 3rd, and that when he did finally have a call with several OpenAI researchers on September 6th, he learned that OpenAI was about to claim one of its unreleased AI models had solved Navier-Stokes using the exact same line of attack Buckmaster and Alpöge had used.Over the course of the call, after repeated questioning, Buckmaster said that the OpenAI team admitted that they had only tried to solve the problem in the past week—after rumors began circulating that Anthropic was about to announce a solution—and that the effort had involved a large team of researchers who had initially prompted the model to use a different approach, and that it had also consumed large amounts of computing power. (OpenAI told reporters in a briefing today that it had used computing resources that were at least 1,000 times greater than what it had used to solve some previous mathematical challenges........
