Generated by Codex with GPT 5.6 Sol XHigh
Techmeme surfaced Terence Tao’s September 11 post, “A Severe Misalignment of AI in Mathematics”. The post publishes an urgent declaration signed initially by 25 Fields Medallists. Its central warning is not that AI cannot do serious mathematics, but nearly the opposite: systems are becoming capable enough that optimizing them to solve famous problems may damage the institutions that turn correct results into shared understanding.
The declaration draws a line between a mathematical answer and the purpose of mathematics. A solved problem can be an important milestone, yet research is not ultimately a leaderboard of true statements. It is a long social process in which people develop concepts, connect new arguments to earlier work, simplify proofs, teach the ideas, and learn how to ask better questions. If AI companies treat open problems mainly as capability benchmarks, they can optimize the most visible proxy while neglecting the intellectual ecosystem that gives a solution meaning.
When the benchmark becomes the goal
Famous problems have traditionally worked as landmarks. Solving one usually signals that a researcher has found a new method or a deeper view of the surrounding mathematical landscape. The result then passes through seminars, criticism, careful exposition, and eventually textbooks. The proof matters, but so does the knowledge created while the community absorbs it.
AI changes the economics and tempo of that cycle. A lab can direct large amounts of compute toward many prestigious problems, announce successes quickly, and use the resulting headlines as evidence that its models have crossed another capability threshold. That can make problem solving an end in itself rather than evidence of understanding. In the declaration’s sharpest formulation, rapidly producing large numbers of correct-or-incorrect claims may exhaust promising questions faster than mathematicians can interpret the answers.
This is a deeper objection than “machines use brute force” or “formal proofs are unreadable.” Even a valid and ingenious machine-generated proof can arrive in a form that does not reveal which ideas are reusable, where the argument sits in the literature, or why its route matters. Correctness is necessary, but mathematics becomes cumulative science only when a community can inspect, explain, and transmit the result.
The benchmark can therefore undermine the thing it is meant to measure. Open problems are useful tests because they stand at the edge of human understanding. If they are consumed as a finite evaluation set, without equal investment in exposition and integration, a lab may demonstrate model capability while reducing the field’s capacity to turn that capability into knowledge.
Mathematics depends on stewardship
The signatories describe students and ideas as the profession’s most valuable resources. Both need time and attention. Students build judgment by struggling with carefully chosen problems, while ideas mature through private discussion, talks, failed attempts, attribution, and revision. These are not ornamental rituals around the production of answers; they are how mathematical taste and understanding are reproduced across generations.
Industrial-scale AI research strains that process in several ways. Results may be announced before specialists can verify them or authors can produce readable accounts. A fast release can obscure the earlier human work that supplied a crucial construction, conjecture, or hint. If machine output incorporates a vast body of prior mathematics without reliable provenance, familiar disputes over citation and plagiarism become harder to resolve. And if human experts must spend increasing amounts of unpaid time checking and rehabilitating machine-generated work, the apparent acceleration may simply transfer the bottleneck from discovery to review.
The declaration’s most important idea is that humans remain essential even when an AI result is correct. A proof does not enter the mathematical canon automatically. Researchers must decide what deserves attention, reconstruct the argument, identify the new concepts, connect it to adjacent work, and teach it. Without that stewardship, AI-generated ideas can remain isolated artifacts rather than living mathematics.
A warning from inside the field, not a rejection of AI
The statement explicitly allows that AI could enhance and accelerate genuine mathematical understanding. It also begins from a strong claim about recent progress: language models can now solve major outstanding problems across several areas. The authors are therefore not defending mathematics by denying the technology’s competence. They are arguing that competence makes governance more urgent.
The breadth of the initial signatories gives the warning unusual weight. The list spans Fields Medal cohorts from 1978 through 2026 and includes researchers from many branches of mathematics. Tao notes that the group released the declaration after only about a week of discussion and without the broader consultation they would normally prefer. That haste is both evidence of perceived urgency and a limitation: the post states principles more clearly than it specifies remedies.
It does not, for example, define an acceptable pace for AI-assisted announcements, a standard for disclosing model and data provenance, or a mechanism for funding independent verification. Nor does it say which problems should be off-limits as benchmarks, how credit should be apportioned among models, labs, prompt authors, and earlier mathematicians, or how students should train when systems can produce advanced results directly. Those omissions do not invalidate the warning, but they identify the institutional work still ahead.
Alignment is also about institutions
In AI debates, alignment often means making a model follow human instructions or values. This declaration describes another layer: the incentives of the organizations deploying AI can be misaligned with the purpose of a profession even when the model produces exactly what was requested. A system asked to solve celebrated problems may succeed, while the surrounding race rewards speed, publicity, and benchmark wins rather than attribution, comprehension, and durable scholarship.
That pattern extends beyond mathematics. Training in science and creative work does more than produce finished outputs; it builds people who can judge evidence, formulate questions, and carry a field forward. When AI can cheaply generate the visible product, institutions need to protect the less visible processes that made the product valuable.
The declaration’s cleanest takeaway is therefore not a demand to stop using AI in mathematics. It is a demand to decide what mathematical progress should mean before capability races decide by default. The relevant measure cannot be only how many open problems fall. It must also include whether new results become understandable, properly attributed knowledge—and whether the next generation still develops the judgment needed to create and care for that knowledge.