Writing archive

· AI for science

AI should make science more ambitious

OpenAI's proposed Navier–Stokes solution and Terence Tao's concerns point to a larger choice: how to turn greater problem-solving capacity into deeper understanding and a more ambitious research agenda.

Deep teal curved planes open into translucent aqua and copper fields around a cream center, with a small cobalt accent.

Reading OpenAI's September 8 announcement about Navier–Stokes, I kept returning to one question: what should researchers attempt when some of the work that once consumed a career becomes possible on a radically shorter timescale?

My answer is that the research agenda should grow. We should investigate harder relationships, connect more disciplines, and attempt experiments that previously demanded more time or expertise than a team could assemble. The ambition should rise with the capability.

Some of the anxiety around this shift feels familiar from the debate about AI in software engineering. A profession develops its identity around difficult work. A new tool starts performing parts of that work. The resulting argument concerns quality, but also what expertise will mean next.

Mathematician Terence Tao's recent comments deserve attention here. He raises a serious objection to treating solved problems as the sole measure of progress. I agree with much of his diagnosis. Where I would push further is in the response: invest in understanding while expanding what science is willing to attempt.

What OpenAI has actually announced

OpenAI reports a proof that a three-dimensional incompressible fluid, initially at rest and subject to a smooth external force, can develop unbounded velocity in finite time while retaining bounded kinetic energy. Its paper, Finite time blowup for Navier–Stokes, states the construction for every positive viscosity.

The forcing matters. This is a claim about the forced equations, not a proof of blowup without an external force. The official Clay formulation explicitly permits smooth forcing in its breakdown alternatives, C and D, which the paper claims to establish. A mathematical singularity in this model also does not mean a physical liquid can acquire infinite speed.

According to OpenAI, the successful group involved roughly 10,000 concurrent agents, with researchers reallocating resources and sharing intermediate insights between groups. The company reports about 88 hours to the result and another 17 for Lean formalization and verification. It used an unreleased internal model, and has published both a paper and Lean artifacts. These are reported research results, not my independent verification of the proof. OpenAI's account of the process.

As of September 9, an announcement should still be distinguished from settled mathematical acceptance. Formal checking must establish the intended theorem under the intended assumptions; the wider community must scrutinize the result. Clay's prize rules separately require publication in a qualifying outlet, at least two years after publication, and general acceptance before consideration.

If the result survives that scrutiny, its significance for this argument is the scale and organization of the research effort. A system that explores many approaches and consolidates partial discoveries changes what a research program can plausibly try.

Tao's concern is about what mathematics learns

In a September 3 thread, Tao distinguishes answering a question from learning through its investigation. A difficult problem can produce useful methods, expose the limits of existing techniques, and reveal connections to other fields. A correct answer can leave much of that value undeveloped.

His Navier–Stokes example that day was hypothetical, preceding OpenAI's announcement. He described a possible outcome in which an autonomous system finds a solution while its operator keeps the discovery process largely inaccessible. The field receives the result, but recovering the useful ideas becomes another substantial research project.

That concern is compatible with enthusiasm for AI. On September 8, Tao welcomed Alpöge and Buckmaster's related work on forced Euler equations, explicitly acknowledging significant AI contributions alongside the authors' effort to make the arguments understandable.

His criticism becomes more contentious when he suggests that some classes of problems may need protection from automated solvers, partly to preserve their value for developing future mathematicians. Later, his September 8 discussion emphasizes a subtler point: infinitely many possible questions do not give us infinitely many fruitful ones. Finding a promising question takes judgment, and rapidly changing, opaque AI capabilities make that judgment harder.

I think that is the strongest version of the objection. Telling researchers to find something harder is inadequate if we offer neither resources nor a way to identify where worthwhile difficulty has moved.

But I would make that identification a research priority. Publish unsuccessful approaches as well as successes. Fund explanation, independent checking, and the search for questions that a new method makes interesting. Preserve opportunities to learn through deliberate training. I would be much more reluctant to make the continued unsolved status of a public research problem the thing we protect.

Engineering offers a useful analogy

The comparison with software engineering is my interpretation, rather than evidence that mathematicians are simply repeating an earlier mistake.

In engineering, I care about whether a system behaves correctly under real constraints. Writing its implementation is part of reaching that outcome. When an agent can handle more implementation work, I want engineers to spend more attention on the requirements, interactions, failure cases, and product possibilities that were previously too expensive to explore.

The corresponding opportunity in science is to change the unit of work. A project might investigate a family of related conjectures instead of a single isolated statement. A laboratory might compare several competing mechanisms instead of committing early to the only one it can afford to study. A team might connect a mathematical model, a simulation, and an experiment within the same research cycle.

These are proposed directions, not capabilities established by the Navier–Stokes announcement. Their value is that they describe what increased capacity could buy beyond a higher publication count.

Researchers will need to adapt, and institutions will need to make adaptation possible. A grant that funds only the first result gives someone little room to develop the explanation. A university cannot ask a junior researcher to compete with industrial compute while providing no comparable access. A training program must still develop the judgment needed to challenge a convincing answer.

Tao's Mathematics in the age of AI makes a related institutional argument: explanation, refereeing, and incorporation into the field's shared understanding deserve greater recognition. I would pair that investment with funding for more ambitious research programs. Understanding and exploration can reinforce each other if we pay for both.

Beyond mathematics, the experiment still matters

We already have examples of AI changing scientific work beyond proof generation.

The AlphaFold 3 paper, published in May 2024, describes a model that predicts structures of complexes containing proteins, nucleic acids, and small molecules. Its authors also identify an important boundary: these are primarily static structural predictions, rather than a full account of molecular dynamics.

The RFdiffusion paper, published in July 2023, goes in a complementary direction: generating new protein designs from molecular specifications. The researchers experimentally characterized hundreds of designed assemblies, metal-binding proteins, and binders. Generation and experimental testing were parts of the same scientific effort.

Those studies support a concrete possibility: we can increasingly use computation to propose candidates for things we want to create, then learn from testing them. They do not establish that a model can solve biology on demand.

My expectation is that much of the next advance will come from tightening that cycle. Generate a candidate, make it, measure it, investigate the mismatch, and use the evidence to choose the next experiment. A failed candidate can still reveal something valuable if the process preserves what failed and why.

Mathematics has unusually strong tools for checking formal deductions. Experimental science must also establish whether the assumptions fit the world. A proof certificate cannot tell a laboratory that a material can be manufactured reliably, or that a proposed biological mechanism is what actually caused an observation.

Greater AI capability should therefore increase our ambition for experimental infrastructure too. Faster proposals create a stronger reason to improve measurement, reproducibility, and the capacity to test competing explanations. Otherwise we merely move the queue from the theorist's desk to the laboratory door.

Scientific ambition depends on trust

The concurrent-work dispute around this announcement also deserves a place in an optimistic account. OpenAI recognizes Alpöge and Buckmaster's priority on forced Euler and says it did not access their work before release. It denies accessing specific user data for the effort, while saying it cannot rule out a contribution from de-identified product-usage data to model improvement. The company's statement.

That statement does not settle every question about credit or trust. My practical conclusion is that researchers need clear terms for unpublished work, meaningful attribution, and access to the material needed to inspect a result. Scientists should be able to use powerful tools without guessing whether sharing a promising direction will undermine their own project.

These conditions help the optimistic future happen. A community that shares methods and partial results can build on a discovery. One that becomes afraid to discuss unfinished work loses part of its ability to collaborate.

We should expect to surprise ourselves

I believe AI will take scientific research to a new level. The strongest reason is the possibility of combining a broader search with more rigorous testing, then using what we learn to formulate questions we could not previously see.

That could eventually let us invent things that even science fiction authors have not imagined. I mean that as a conviction about the potential of discovery, not a forecast established by one mathematical result. Naming the inventions now would miss the point: some will depend on concepts and connections that we have yet to discover.

For a research leader, the immediate question is practical. Which valuable investigation did your team rule out because it required too many calculations, too many disciplines, or too many iterations? What would have to change for you to attempt it, and what evidence would convince you that it worked?

Keep the rigor. Give researchers the tools, time, and institutional support to aim higher. The next research agenda should be larger than the one these systems inherited.

Sources