Generated by Codex with GPT 5.6 Sol XHigh
Techmeme surfaced Fields Medalist Timothy Gowers’s August 12 essay, “What sort of maths are LLMs good at?”. Written days after OpenAI announced ten claimed advances in mathematics and theoretical computer science, the essay asks a more useful question than whether language models are simply “good at maths”: what kinds of mathematical search currently play to their strengths, and where might human judgment still matter?
Gowers starts from a striking pattern. The most celebrated recent AI results include an explicit non-sofic group, a superexponential lower bound for multicolour Ramsey numbers, and work related to the Jacobian and unit-distance conjectures. Models can prove difficult universal statements, but many of their strongest headline results are examples or counterexamples—objects that establish existence or overturn an expectation—rather than proofs built around a surprising new conceptual argument.
That observation does not lead to the easy conclusion that models are inherently better at existential statements. Gowers instead develops a sharper hypothesis: current systems combine broad mathematical recall with the ability to explore far more candidate paths than a person can. That makes them unusually effective when progress comes from trying many plausible constructions, standard techniques, or variations until one works. The possible human advantage is not raw mathematical knowledge or speed, but a researcher’s “nose” for pruning a huge search tree before wasted branches multiply.
Not just a quantifier trick
Mathematical statements can often be rewritten so that a theorem looks like an existence claim or an existence claim looks universal. A result saying that every large integer is the sum of three primes can be expressed using alternating existential and universal quantifiers, but no mathematician would naturally describe it as finding a counterexample. Conversely, a construction that supplies widely separated normed spaces in every dimension is normally understood as producing examples, even though its formal logical structure looks similar.
The category also depends on mathematical context. An object feels like a counterexample when it defeats a claim researchers had reason to believe. The same object may later be described simply as the first example of a newly recognized class once confidence in the old conjecture has faded. Gowers notes that the first non-sofic group arguably belongs in this second category: experts had proposed several routes to such a group, so the result supplied a long-sought object more than it shattered a deeply held belief.
This matters because syntax alone cannot explain model performance. Existential steps are everywhere inside ordinary proofs: a researcher may need to find a stronger induction hypothesis, an intermediate property, an invariant, or a useful construction. If AI systems have a genuine advantage on certain examples, the cause must lie in how those examples are found rather than in whether the finished theorem begins with “there exists.”
Search breadth versus mathematical taste
Gowers sketches several ways mathematicians search for examples. They test standard candidates, combine familiar constructions, leave parameters unspecified until later constraints determine them, attempt the opposite theorem to expose a weak point, refine failed guesses, build an object stage by stage, or use random and generic choices. Some of these methods reward exactly what a frontier model can supply: a large library of known techniques and the speed to test many low-probability options.
This offers a plausible explanation for expert reactions to AI-generated mathematics. A result may initially look astonishing because a famous problem has fallen, yet its final approach can seem recognizable once exposed—something an expert might have found with one well-chosen hint. A model need not invent an entirely new mathematical language if it can search a vast space of existing ideas cheaply enough to discover the rare combination that works.
Humans may still be better at problems whose search trees are both deep and heavily branched. Experienced researchers often abandon an approach long before they can formally prove it will fail; they sense that a lemma is becoming irrelevant, that a reduction is circular, or that a promising-looking direction is not producing genuine leverage. Gowers sees current models repeatedly offer approaches that sound good until examined closely, or announce that they have reduced a problem to a narrower question several times without moving nearer to a solution. The weakness is not an inability to generate steps. It is difficulty distinguishing progress from motion.
Brute-force strength may even conceal that weakness. Because models can search faster and draw on more recalled mathematics, an inefficient process can succeed before combinatorial explosion becomes visible. Scaling would then extend the range of problems the process can reach without necessarily teaching the model to prune like a strong researcher.
A better test of mathematical intelligence
The training record makes this boundary hard to measure. Published papers present polished proofs and usually erase abandoned approaches, failed constructions, and the judgments that redirected the author. Models learn from the destination while seeing little of the navigation. When a system produces a deep-looking idea, observers also cannot easily tell whether it reasoned its way there or reconstructed an obscure argument already present in its training data.
Gowers proposes two useful directions. Researchers could test somewhat weaker models that have been shielded from relevant literature on carefully designed problems, then ask whether the solved cases share a recognizable search structure. Training could also reward reaching a correct result while penalizing excessive dead ends or reliance on a known answer. That would test whether systems can learn not only to traverse mathematical search trees, but to judge which branches deserve attention.
His standard for a decisive advance is revealing: a model should produce a proof as conceptually surprising as the 2016 solution to the cap-set problem, whose method sharply improved previous bounds and opened an unexpected line of research. Such a result would demonstrate more than persistence across familiar techniques. It would show the ability to find a path that expert humans did not even know belonged in the tree.
The essay’s durable insight is that the human–AI boundary in mathematics is unlikely to be “proofs versus counterexamples.” It is closer to the geometry and economics of search. Models already have extraordinary breadth and can spend computation on paths people would ignore. Humans may retain an edge where success depends on ruthless, intuitive pruning and on recognizing a conceptual route before there is much evidence that it will work. Given the pace of change, Gowers treats that edge as a present hypothesis rather than a permanent refuge—but it is a much more precise hypothesis than saying machines merely lack creativity.