OpenAI announced that it has solved the Navier-Stokes problem, which mathematicians have been working on for nearly 90 years. The problem, which concerns the equations that describe the motion of liquids and gases, is among the seven Millennium Prize Problems identified by the Clay Mathematics Institute, with its verified solution qualifying for the $1 million prize. According to the development also reported by The New York Times and Wired, the company achieved the solution by taking advantage of 10 thousand agents working simultaneously with an internal artificial intelligence model that is stronger than the GPT-6 Astra it recently launched. While OpenAI’s statement points to a remarkable result in terms of mathematics, a serious data usage debate has begun among researchers about how the solution emerged. At the center of the debate is the work carried out by New York University mathematics professor Tristan Buckmaster and Anthropic researcher Levent Alpöge around the same problem.
According to the information provided by OpenAI, training of the internal model in question started on August 28. The company states that the model, which is not yet available for public use, has demonstrated unprecedented performance in its own evaluations, especially in the field of mathematics. Navier-Stokes equations are used to mathematically model the movement of fluids and cover a wide range of physical fields, from air currents to the movement of water. The millennium problem, on the other hand, is about basic mathematical questions such as whether uniform solutions of these equations in the three-dimensional case exist under all conditions and whether these solutions develop singularities over time. Therefore, before OpenAI’s work can truly be considered a solution to the Millennium Problem, it requires detailed review and verification by the mathematical community beyond the company’s statement.
Data usage discussed in OpenAI Navier-Stokes solution
The notable aspect of the discussion was that just a day before OpenAI’s announcement, Buckmaster, together with Alpöge, a researcher at Anthropic, published results on a related problem. According to Buckmaster, the researchers contacted the company after learning that OpenAI was aware of the progress of their work. It was then revealed that OpenAI’s proof for the Navier-Stokes equation followed a path related to the approach Buckmaster and Alpöge were working on. The fact that the two researchers benefited from OpenAI Codex and Anthropic Claude throughout the project made the discussion more than just a matter of mathematical similarity. Buckmaster stated that they entered all project drafts into the Codex sessions and raised the possibility that this data may have reached OpenAI’s internal model.
Buckmaster says that in his meeting with the company, he directly asked whether the model was accessing user data on Codex. Stating that the answer to his question was whether the model had searched for user data, the mathematician stated that he did not receive an answer when he asked again whether the same data was used in model training. OpenAI said in a statement on Tuesday that data belonging to a specific user was not used to solve the problem. However, the company acknowledged that it could not completely exclude the possibility that de-identified data derived from the use of the products may have contributed to the development of the models. This distinction constitutes one of the main points of the discussion; Because directly accessing a user’s work sessions and the data derived from these sessions later being included in model training mean different processes in terms of technical and data governance.
Sebastien Bubeck, from the OpenAI technical team, also stated that they did not see Buckmaster and Alpöge’s work before it became public. According to Bubeck, the evidence of the two parties differs significantly when compared retrospectively, and the definitive conclusions proven are not the same. Buckmaster did not find this explanation sufficient in his response via Mastodon. The researcher believes that OpenAI’s statements clearly acknowledge that training data from a period after the date they obtained their results may have been used. The dispute thus expands not so much on whether OpenAI directly accessed researchers’ unpublished work, but rather on the potential impact of data derived from users’ research sessions with AI tools on subsequent models.
This situation makes visible another question that arises with the increasingly intensive use of artificial intelligence tools in scientific research. When researchers import unpublished proofs, code, and working drafts into systems such as Codex or Claude, the conditions under which this content is stored and whether it is used in the development of future models can have direct consequences for scientific priority. It is possible that a model’s educational background can make it difficult to identify the source of research ideas, especially when independent teams address the same mathematical problem. However, it is a possibility that cannot be ignored that similar mathematical methods can be discovered independently. In order to clarify the disagreement between OpenAI and Buckmaster’s statements, it is therefore necessary not only to compare the proofs, but also to sufficiently clarify the training process of the relevant model and the use of data.
OpenAI also announced that it does not plan to receive the $1 million Millennium Prize set for the Navier-Stokes solution. However, giving up the reward does not change the need to independently evaluate the mathematical result presented. If the company’s work is confirmed, it will provide a strong example that artificial intelligence systems can be used more comprehensively in advanced mathematics than auxiliary tools that merely process existing information. In contrast, the questions raised by Buckmaster and Alpöge show that in evaluating such success, the source and conditions of use of the research data must be examined as much as the model performance. At this stage, considering OpenAI’s mathematical claim and researchers’ objections to data use separately provides a more balanced picture in terms of both the scientific quality of the solution and the working conditions of AI-supported research.
Join Channel