Tim Gowers analyzes what mathematics LLMs excel at after OpenAI's breakthroughs
Fields medalist examines whether LLMs are better at finding counterexamples than proofs after OpenAI solved ten major problems.
Conversation activity · last 13 hours peak 5/30m
Latest coverage newest 5 of 5 items
Summary, timeline and people extracted by Claude from 26 items across 4 sources · 2h ago. Quotes are verbatim.
What to know
- OpenAI recently solved ten major mathematics problems including a non-sofic group construction and a Ramsey theory result, but Gowers argues LLMs are not yet better than humans at all mathematical tasks.
- Gowers suggests LLMs may be particularly well-suited to finding counterexamples, but notes this pattern isn't universal and remains unexplained theoretically.
- The conversation reflects broader uncertainty: commenters debate whether results reflect sampling-based search capabilities, test-time scaling effects, or genuine mathematical insight; cost and verification bottlenecks are practical concerns.
How it unfolded
-
A commenter raises concerns about whether the pace of results would be faster if LLMs truly reached human level, noting a report of progress on the Riemann hypothesis and questioning who can actually verify such proofs.
“How many people in the world can actually verify a proof? How many would be willing to dedicate the time required?”
parhamn · Hacker News ↗ -
Commenters discuss whether the results reflect LLM capabilities for counterexample-finding or represent broader implications of test-time scaling, with some arguing sampling is fundamentally what AI excels at.
“Sampling is what AI is good at. Making examples and doing LeetCode are similar in that verification is clear and cheap.”
h_mirin · Hacker News ↗ -
Tim Gowers posted a detailed blog post examining what kinds of mathematics problems LLMs are good at, noting that while the OpenAI results are extraordinary, LLMs do not appear better than humans at all aspects of mathematics.
“These results, and the other eight on the list, are extraordinarily impressive, but it still doesn't seem to be the case that LLMs are better than all humans at all aspects of mathematics.”
Tim Gowers · Hacker News ↗ -
OpenAI announced it had solved ten major problems in mathematics and theoretical computer science, including the first construction of a non-sofic group and proof that the multicolour Ramsey number grows superexponentially.
What people are saying verbatim
“These results, and the other eight on the list, are extraordinarily impressive, but it still doesn't seem to be the case that LLMs are better than all humans at all aspects of mathematics.”
Tim Gowers, Fields medalist mathematician · Gowers's Weblog ↗ · Aug 11
“A good sign that LLMs have reached human level for a much wider class of problems will be if they start proving theorems using methods that, like much of the very best human mathematics, are new and surprising but that with hindsight come to seem beautiful and natural.”
Tim Gowers, Fields medalist mathematician · Gowers's Weblog ↗
“Sampling is what AI is good at. Making examples and doing LeetCode are similar in that verification is clear and cheap.”
h_mirin, Hacker News commenter · Hacker News ↗ · Aug 12, 7:02 AM
“It seems intuitive that finding a counter-example might be easier than proving a generality, since you're starting from a concrete goal that you can branch out from, identify sub-problems, etc.”
HarHarVeryFunny, Hacker News commenter · Hacker News ↗ · Aug 12, 10:46 AM
“Increasingly, I've begun to think of LLMs as sources of really interesting random objects: large pieces of "reasonable thinking" conditioned on a task.”
tel, Hacker News commenter · Hacker News ↗ · Aug 12, 8:30 AM
The conversation positions from the crowd, verbatim
The discussion reflects measured but contested optimism about LLM capabilities. Commenters acknowledge the breakthrough results as impressive but scrutinize the mechanisms behind them—sampling vs. reasoning, cost-effectiveness, and verification limits—while remaining uncertain whether LLMs will achieve genuine mathematical creativity or remain bounded by search-like behavior.
LLMs excel at sampling and search rather than reasoning; they're good at counterexample-finding because verification is cheap and the search space is concrete.
-
“Sampling is what AI is good at. Making examples and doing LeetCode are similar in that verification is clear and cheap.”
h_mirin · Hacker News ↗ -
“It seems intuitive that finding a counter-example might be easier than proving a generality, since you're starting from a concrete goal that you can branch out from.”
HarHarVeryFunny · Hacker News ↗ -
“Increasingly, I've begun to think of LLMs as sources of really interesting random objects: large pieces of "reasonable thinking" conditioned on a task.”
tel · Hacker News ↗
Practical concerns about cost, verification bottlenecks, and human oversight limit the real-world impact of these results.
-
“How many people in the world can actually verify a proof? How many would be willing to dedicate the time required?”
parhamn · Hacker News ↗ -
“It is often overlooked how expensive these models can be to run, and the false positives. You are going to be burning through a lot of $ if you use the latest models on hard problems, with no assurance of progress...”
paulpauper · Hacker News ↗
LLMs lack the creative insight that characterizes the best human mathematics; true breakthrough would require proving theorems via novel, surprising methods.
-
“'generative' ai is not good at generative science, of which requires unique human perception that is not purely symbollic manipulation, but requires a form of revelation.”
justanotherjoe · Hacker News ↗
Voices from the web unedited
-
> A good sign that LLMs have reached human level for a much wider class of problems will be if they start proving theorems using methods that, like much of the very best human mathematics, are new and surprising but that with hindsight come to seem beautiful and natural. They should also be methods that are difficult to stumble on by accident. It…
-
🌘 大型語言模型擅長哪類數學? | Gowers 的網誌 ➤ 探討 AI 攻克高等數學難題背後的邏輯與極限 ✤ https:// gowers.wordpress.com/2026/08/1 2/what-sort-of-maths-are-llms-good-at/ 本文探討大型語言模型(LLM)在數學領域的當前能力與侷限,特別針對 OpenAI 近期宣佈解決多項重大數學難題的背景進行分析。作者指出,LLM 在尋找數學「反例」上展現出非凡的優勢,並透過解析量詞結構與巴拿赫-馬祖爾幾何學中的格盧斯金定理,深入探討為何特定類型的數學命題更適合 LLM 來求解,為評估人工智慧在高等數學研究中的定位提供了系統性的思考框架。 +…
-
> If they were, then their big speed advantage over us would mean that there would be much more of a flood of results.Is this true right now? Just recently Jarred Sumner tweeted [1] that he managed to make some progress on the Riemann hypothesis while on a jog. Managed to get somewhere by encouraging the llm to “keep going” and “believe in…
-
This is really an argument about test-time scaling, even though the post never uses the term.These days "test-time scaling" mostly means letting the model talk to itself for longer, but the first genuinely surprising results came from plain sampling. Google's AlphaCode generated millions of candidate programs and filtered them down to a handful of…
-
For a list of AI accomplishments in mathematics see https://mathoverflow.net/questions/502120/examples-for-the-u... - or a candidate list here: https://aimath.robertj1.com/ . Many have observed an affinity of AI to the search for counterexamples - or examples. Looking at afore lists, something much more sociological crosses my mind: There is a…
-
Increasingly, I've begun to think of LLMs as sources of really interesting random objects: large pieces of "reasonable thinking" conditioned on a task. It's not that these are correct, in general, but instead they're a concentrated form of random search where that "randomness" is very likely to follow plausible, human patterns.You can toss it at a…
-
Given coding agent's demonstrated difficulties with concurrent code, even relatively simple concurrent code, it would be interesting to see how they do with temporal logic. I don't know enough to throw AI at the problems in that space but I wonder if they wouldn't crash and burn on it.(I haven't had the opportunity to throw a current-gen frontier…
-
A thoughtful and measured post, as usual from Gowers. The final note is neat and worth pasting out here in full:> A good sign that LLMs have reached human level for a much wider class of problems will be if they start proving theorems using methods that, like much of the very best human mathematics, are new and surprising but that with hindsight…
-
Correct me if I'm wrong, this is not the right way to ask this question.LLMs are good at pattern recognition, so its less a type of math that they'll be good at, and more that when you provide documentation or text that can be easily parsed/compared to its training data/reasoning ability, the better answers you get from an LLM.Also, you need to be…
-
It seems intuitive that finding a counter-example might be easier than proving a generality, since you're starting from a concrete goal ("build a foo that has properties X, Y & Z") that you can branch out from, identify sub-problems, etc.Proving a generality seems much more difficult since you don't know what you are trying to build, although I…