194 pointsby ColinWrightAug 12, 2026

18 Comments

n4r9Aug 12, 2026
A thoughtful and measured post, as usual from Gowers. The final note is neat and worth pasting out here in full:

> A good sign that LLMs have reached human level for a much wider class of problems will be if they start proving theorems using methods that, like much of the very best human mathematics, are new and surprising but that with hindsight come to seem beautiful and natural. They should also be methods that are difficult to stumble on by accident. It is hard to say precisely what would count as such a proof, but I think we’ll recognise it when we see it.

tcp_handshakerAug 12, 2026
>>A good sign that LLMs have reached human level for a much wider class >> of problems will be if they start proving theorems using methods that, like much of the very best human mathematics, are new and surprising but that with hindsight come to seem beautiful and natural.

I must be taking crazy pills and the AGI surely will pass me by... But TODAY, middle August 2026...And in the context of testing and evaluating the capabilities of current SOTA models to implement an Agentic application for job search, here is some simple inhouse built evals I run today, since I don´t trust LLM vendors published benchmarks...

Models tested: GPT-5.6 Sol in Extra High mode and Opus 4.8 Max.

TASK REQUEST: Clear, not too long not too short prompt, for LLMs to go out and research freelance consulting gigs for one specific IT domain, and in one specific country in Europe, including maybe opportunities driven from temp agencies based in geographically close countries.

RESULT: Models go out, fetch the data, and completely misunderstand the task...offering on first results, permanent roles instead of freelance, and based on the country where the agencies are, not in the one it was request for. Think for example IT jobs in Ireland, while freelance agency in London.

ANALYSIS: No intelligence I can call it shown by models, adding cognitive effort for human in the loop to detect subtle factors, and therefore totally useless for agentic app...Best practices would be I guess to add agents on top of agents but although in the p95 of cases that will reduce the errors...for the remaining 5% that could have hallucinations or logic hallucinations like these ones, compounding on top of other logic hallucinations.

I dont care about the theorems being proven. At the end we will found out what most mathematicians were doing, was just exploring the same combinatorial and abstraction patterns. And because of that I am sure LLMs will make mince meat of a lot of mathematical domains.

But right now, what we call intelligence is not existing where it matters, and Ed Zitron is right its a parlour trick.

m348e912Aug 12, 2026
>TASK REQUEST: Clear, not too long not too short prompt, for LLMs to go out and research freelance consulting gigs for one specific IT domain, and in one specific country in Europe, including maybe opportunities driven from temp agencies based in geographically close countries.

As a human, not an LLM, I could interpret "including maybe opportunities driven from temp agencies based in geographically close countries" as meaning "including opportunities in nearby countries outside of Ireland" (that happen to be driven by temp agencies).

Before writing off LLM as simply a "stochastic parrot" or a "parlour trick" remember it can't read your mind, not yet anyway.

tcp_handshakerAug 12, 2026
I am describing the contents of the prompt that was not the prompt. The prompt was very clear to the model that some freelance opportunities in country A, the only one in consideration could be available via agencies in country B and C. And it was a clear prompt.

So what happen is a prompt said for example, find freelance opportunities in Ireland but keep in mind some of these might be available via temp agencies in London.

If you offer me not freelance but permanent roles, and not in Ireland in London...that is a logic failure.

Its this type of complexity with the normal world, that these SOTA constructions so badly fail at, and so spectacularly fail at the margins... despite maxing all benchmarks...Parlour trick.

someguyiguessAug 12, 2026
Based on the rest of your writing I’m going to assume that the prompt was the problem.
coldteaAug 12, 2026
He was perfectly clear in both cases.

If a human misunderstood this, they'd be a dumb human.

fn-moteAug 12, 2026
For perspective, I agree with the GP. The writing is not perfectly clear. We don’t have enough evidence to know if that was part of the problem.
tcp_handshakerAug 12, 2026
Keep deluding yourself, unless you work for an LLM provider...

"Frontier LLMs Still Struggle with Simple Reasoning Tasks" https://arxiv.org/abs/2507.07313

"General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks" https://arxiv.org/abs/2604.11778

"...General365, a benchmark specifically designed to assess general reasoning in LLMs. By restricting background knowledge to a K-12 level, General365 explicitly decouples reasoning from specialized expertise. The benchmark comprises 365 seed problems and 1,095 variant problems across eight categories, ensuring both high difficulty and diversity. Evaluations across 26 leading LLMs reveal that even the top-performing model achieves only 62.8% accuracy, in stark contrast to the near-perfect performances of LLMs in math and physics benchmarks..."

xnorswapAug 12, 2026
I found it very difficult to parse your description, "Clear, not too long not too short prompt, for LLMs to go out and research freelance consulting gigs for one specific IT domain, and in one specific country in Europe, including maybe opportunities driven from temp agencies based in geographically close countries."

( I was trying to quote a single sentence and then realised it ran on for the whole paragraph. )

Given how difficult I found that to follow, are you sure your prompt is actually "Clear, not too long not too short"? We now only have your word for it. I too had assumed that was a prompt given to an LLM to further prompt agents.

geonAug 12, 2026
That was not the prompt.
xnorswapAug 12, 2026
I know that. We don't know what the prompt was. We only have a self-assessment of the quality of the prompt from the person who wrote it.

It sounds like they're hitting a data source quality issue, which is hardly uncommon in scraping.

It's common for job boards to obscure who the real clients are, and if the scraping engine is LLM powered ( rather than LLM written ), then I would expect it to accidentally present agencies as the contracting organisation sometimes.

Breaking down the process so you can inspect the messy middle of a data pipeline is an important part of software engineering, but it sounds like they've tossed a messy task at an LLM and expected it to be proficient end-to-end.

tcp_handshakerAug 12, 2026
You can easily test this yourself with the SOTA models....or read the corroborating literature...

"General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks" https://arxiv.org/abs/2604.11778

"...General365, a benchmark specifically designed to assess general reasoning in LLMs. By restricting background knowledge to a K-12 level, General365 explicitly decouples reasoning from specialized expertise. The benchmark comprises 365 seed problems and 1,095 variant problems across eight categories, ensuring both high difficulty and diversity. Evaluations across 26 leading LLMs reveal that even the top-performing model achieves only 62.8% accuracy, in stark contrast to the near-perfect performances of LLMs in math and physics benchmarks..."

xnorswapAug 12, 2026
What proportion of the human population could answer the example from that paper?

    Question: Strangers A, B, C, D, and E line up from youngest on the left to oldest on the right. Their clothing
    colors and shoe colors all differ, and they come from five different regions.
    Known facts:
    1. A is from Morocco.
    2. D is five years older than B.
    3. E is older than A.
    4. C stands next to D.
    5. A stands next to B.
    6. The person in teal shoes is not adjacent to the person from Vanuatu.
    7. One twelve-year-old wears yellow shoes.
    8. The person in orange shoes wears white clothing.
    9. The person in blue clothing is from Chile.
    10. The youngest person wears red shoes.
    11. Counting from the right, the fourth person comes from South Africa.
    12. E wears yellow clothing.
    13. The person in green shoes does not wear multicolored clothing.
    14. Two people are twelve years old, ordered by birth month.
    15. One adult is thirty-five years old, and that age is sixteen less than the combined ages of the other four.
    If you multiply every possible age C might have, what number do you obtain?

What does "The person in green shoes does not wear multicolored clothing" even mean?

Nowhere is "multicoloured" defined, are we to assume it should be treated as a colour and implied that someone else must be wearing "multicoloured clothing"? Because strictly that doesn't logically follow, and it ought to be phrased as "The person in green shoes is not the person wearing multicoloured clothing" if that is the case.

This is an extremely hard logic puzzle, especially since it's revealed at the end that there are multiple solutions.

I'd expect anyone to struggle unless armed with prolog.

defrostAug 12, 2026
> What does [ 13 ] even mean?

It means that it is possible for someone to be wearing a white shirt and yellow pants (say), but the person in green shoes came from the set of a Wes Anderson film.

xnorswapAug 12, 2026
That's cute, but it's driving me crazy, I guess I'll have to sit down and solve it to figure out how it's meant to be clued, assuming it's not just a red-herring entirely.
xnorswapAug 12, 2026
Right, this is based on pen-and-paper working, so I could be wrong, but I think it's a red-herring, the set of "clothing" seems to be:

white, blue, yellow, and is otherwise undefined.

But we know from [1], [2], [4], [5] and [11], that the order must be:

A, B, C, D, E or A, B, E, D, C.

Which makes C either the older 12 year old or the 35 year old.

The key this is that D can't be a 12 year old without A being in slot 2, but A can't be in slot 2 because slot 2 is South Africa and A is Morocco.

Trying to reach the shoes + Vanuatu clue is a complete waste of time, paying any attention to shoes or clothing is a waste of time, it feels like there ought to be a way to narrow it down to one of those two configurations, but the clothing is too ambiguous, the shoes end up irrelevant.

What a frustrating puzzle, where half the clues are seemingly redundant.

retsibsiAug 12, 2026
I independently reached the same conclusion, and I think we must be right, because the thing I thought I might have missed was some way in which the shoes/clothes are actually relevant and rule out one of those orderings. But our answer matches the one in the paper, and I don't see a mistake we could have made to reach the right answer in the wrong way.

I guess they were not optimising for a satisfying solving experience! (I'm not sure what 'expansion' means in this context, but it sounds like maybe the 'expansion strategy' referred to in the paper involved padding out problems with red herring premises?)

fn-moteAug 12, 2026
Super interesting, thanks for the reference.

I particularly found the note about “local collapse” helpful (near the end of section 3). The idea is that even though benchmarks contain a wide variety of different reasoning tasks, each individual problem requires only a few skills - unlike this benchmark where they deliberately construct tasks that span many categories.

geonAug 12, 2026
Would the llm work better if it was given the job ad and asked where the job was located?

It seems to me that such simplified tasks tend work better. The rest of the loop is just scraping websites, which doesn’t really have a reason to rely on ai agents.

CodeCompostAug 12, 2026
LLMs have no sense of geography. They measure distances between parts of words, not distances between parts of world.
foco_tubiAug 12, 2026
“Canada is north of the United States” has a higher probability of correctness than “Canada is east of Greenland.”
fragmedeAug 12, 2026
> TASK REQUEST: Clear, not too long not too short prompt

Why do you think there's such a thing as too long for an LLM prompt? You'll run into context window limits at some point, but the more verbose you are with what you ask of it, the better the results will be.

root-parentAug 12, 2026
>> t the more verbose you are with what you ask of it, the better the results will be.

Trivially falsifiable:

"Context Length Alone Hurts LLM Performance Despite Perfect Retrieval"

https://aclanthology.org/2025.findings-emnlp.1264/

"Large Language Models Can Be Easily Distracted by Irrelevant Context"

https://arxiv.org/abs/2302.00093

simdezimonAug 12, 2026
https://claude.ai/share/f0c5c3c9-1882-44b1-b30e-fb427a0df472

I don't know your prompt and setup, but my claude had no problems doing that task. The search index isn't live, so it can't find current gigs, but that is a tooling problem.

vhantzAug 12, 2026
Stop wasting your time and use actual code for most of what you give an LLM to do. Make them write the code even.

Anything that can be verified mechanically should be code. Only use LLMs to fill in the gaps where things are fuzzy. Don't fall for the idea that those harnesses are general purpose, make your own fit to your task with the guards and verification steps you need. Make the LLM create the harness even.

There is no amount of markdown that can make a machine generating plausible text generate truthful text, it just happens to be truthful because of what it was trained on. Nothing coming out of an LLM should be taken at face value.

The propaganda about LLMs being intelligent and able to "reason" is only serving the companies selling you tokens to waste on "prompt engineering".

root-parentAug 12, 2026
The irony about your comment is that, this is probably the most likely opinion and most consensual around many technological practitioners.

But the mathematicians here in this thread, are having a hard time with these clearly dumb models, doing so well in proving theorems in their domains :-)

groundzeros2015Aug 12, 2026
This is an off topic rant unrelated to mathematical ability which is a closed problem often with complete logical information.
Anon1096Aug 12, 2026
Reading the responses to your comment the discussion would be a lot more productive if you shared your logs (preferably several of different top models since that's what you're claiming) where LLMs fail at this. Not very useful for people to go back and forth speculating on what you could have asked and with what formulation. As it stands for me simdezimon's logs are pretty definitive that there shouldn't be any problem for current capabilities agents to solve the task.
hansvmAug 12, 2026
I'm not sure why exactly, but I've heard the same from every single person using LLMs for anything related to jobs. The posting says it needs at least a B.S., and the LLM denies an applicant because they have an M.S. The posting thinks it needs 3yrs of work experience in XYZ technology, and it won't add it to the candidate's list because it doesn't have the context that the HR/LLM filter on the posting adds a bunch of nonsensical requests or that some other combination of skills makes the candidate stand out above and beyond that missing "requirement." And so on. The quality is quite poor.

On the other end of it, something like 80% of resumes I receive right now are clearly hallucinated -- referencing accomplishments that are copy-pasted from the novel-to-our-company thing in the job description a candidate will be working on, usually claiming they did XYZ at big tech a decade before the thing existed, or similarly with languages and skills. The resume "tailoring" process just manufactures lies rather than tailoring actual experience to the actual job.

pinkmoonxAug 12, 2026
How interesting is it that in the same way the human brain unconsciously does calculus and linear algebra, but struggles in the conscious space (we have to go learn it, it’s not easy) the same is true of LLMs.

They are algebra, and yet kinda suck at it without training

philipwhiukAug 12, 2026
> the human brain unconsciously does calculus and linear algebra

We're not unconciously doing calculus and linear algebra. If you're arguing you 'do calculus' to predict how to catch a ball, I'm sorry but it's not supported by the data.

greenmoonxAug 12, 2026
King of nitpicks, yes ofc not literally and it's the whole point of what was said.

Imagine three neurons, each with a firing speed at a time called τ: a₁(τ), a₂(τ), and a₃(τ). Together, these three firing speeds make a group: A(τ) = [a₁(τ), a₂(τ), a₃(τ)]. This group shows how active the three neurons are at that moment. You can think of this group as a point in space with three directions, one for each neuron’s activity. This helps us see what the brain is doing when a ball flies through the air.

It doesn't mean the brain is articulating the language of math behind the scenes.

You simply remember (record/store) values of where the ball was last time you saw it fly through the air. You get better at modeling the trajectory because the neurons physically move closer together as you learn. We can and do represent this with math.

These linear algebra vectors and algorithms are math/compsci that can represent the branching nature of firing neurons in the same way it can be applied to how a river winds through a landscape, and other things. (this does not mean "the river is doing math" btw).

That's why in neuroscience, linear algebra is so widely used to represent firing rates of neurons, coordinate systems in sensory spaces, and multi-channel neuroimaging data matrices (like fMRI or EEG). Also PCA (dimensionality reduction) relies solely on it since you are representing actual brain cells with elements in Arrays.

> not supported by the data

Entire fields exist, you're just ignorant beyond belief.

thinkharderdevAug 12, 2026
> You get better at modeling the trajectory because the neurons physically move closer together as you learn

I think "modeling the trajectory" is not necessarily what we are doing either. It's more likely we are using much simpler heuristics. If you are trying to catch a ball flying through the air, you can just look at the ball and modulate your running speed to keep your eyes at a fixed angle until you catch the ball. It's much more analogous to a PID controller than a model of the trajectory.

h_mirinAug 12, 2026
This is really an argument about test-time scaling, even though the post never uses the term.

These days "test-time scaling" mostly means letting the model talk to itself for longer, but the first genuinely surprising results came from plain sampling. Google's AlphaCode generated millions of candidate programs and filtered them down to a handful of submissions, which beat the average human programmer in 2022, before ChatGPT even showed up.

Sampling is what AI is good at. Making examples and doing LeetCode are similar in that verification is clear and cheap. Compared to that, "proof" is still a vague concept, except where Lean works. See the fuss over the ABC conjecture. So humans are still needed.

The interesting question to me is what happens after enough learning from "sampling." Isn't AlphaGo's move 37 an AI's nose? If that happens in mathematics, we may end up with results that are correct, machine checkable, and not explainable in any way we find satisfying.

laszlojamfAug 12, 2026
for somebody who's out of the loop: what's the fuss over the ABC conjecture?
pringk02Aug 12, 2026
https://en.wikipedia.org/wiki/Inter-universal_Teichm%C3%BCll...

Wikipedia is maybe the narrow end of a wedge into this topic but the controversy revolves around a very large and very complex paper that few people are equipped to understand and some of those who are able believe the proof is false.

steinwindeAug 12, 2026
This is a reference to Inter-Universal Teichmüller Theory. Its Wikipedia article gives a good overview (https://en.wikipedia.org/wiki/Inter-universal_Teichm%C3%BCll...). In maths lasting disagreements over a published "proof" are rare, but IUTT is an example of it. What the article misses: There is a more recent, ongoing effort to formalize the published proof in Lean under the name of "LANA" (e.g. see https://zen.ac.jp/news/zmcpostevent0717e and https://github.com/katobungen/LANA_report_202607/blob/pdf/LA... for a recent update). I guess most mathematicians agree that a successful compile of the proof in Lean would confirm its validity. My personal impression is that the process got stuck at the very point Peter Scholze and Jakob Stix pointed out 8 years ago. Officially LANA has still not reached a conclusion.
brazzyAug 12, 2026
TDLR for the other two comments: a Japanese mathematician is claiming to have a proof for it, but it is based on an entirely new very complex field of maths which he invented. Getting into it takes years, so other mathematicians are hesitant to invest that much time only to find out that the proof is broken and the field isn't otherwise useful.

It doesn't help that the author is rather withdrawn and not willing to spend any effort in making it more approachable.

Some tried, and said they found gaps in the proof, to which the author responded, but they were not convinced.

And that's essentially the situation since 2018.

jgalt212Aug 12, 2026
> Google's AlphaCode generated millions of candidate programs

The trick is avoiding the infinite monkey problem. If your problem is amenable to RL, then you probably don't even need an LLM, Monte Carlo Tree Search gets you there with less expensive hardware.

scronkfinkleAug 12, 2026
> A good sign that LLMs have reached human level for a much wider class of problems will be if they start proving theorems using methods that, like much of the very best human mathematics, are new and surprising but that with hindsight come to seem beautiful and natural. They should also be methods that are difficult to stumble on by accident. It is hard to say precisely what would count as such a proof, but I think we’ll recognise it when we see it

Agreed. I find that after seeing these results from OpenAI we undeniably have a machine that has:

* General knowledge of nearly every subject humanity has ever learned

* The ability to simulate reasoning (albeit sometimes not very well) with that knowledge

* The ability to reference across the domains of knowledge

To me, this is more or less what I would think "Artificial General Intelligence" is. It's the cumulative knowledge of all general human intelligence, baked into an artificial form, which can then use that knowledge to achieve novel goals.

In many cases of mathematical breakthroughs there is an insight that comes from just happening to know a combination of already existing ideas and then combining them to solve that problem. This is where having that general knowledge seems particularly strong because we can run these machines for weeks on end effectively trying to brute force.

That being said, I could never imagine an LLM in its current form inventing something as elegant as the Fourier transform.

porridgeraisinAug 12, 2026
I share the same thinking. What do you think is a good way to try to define this "elegance"? If we try to use the mental framework of

Step 1. LLM "brute forces" a search

Step 2. We train on this trace

Step 3. In the next model, LLM internally makes a "shortcut" for this path and "brute forces" it quicker (or one shots its in the best case)

And we want to ultimately show why that definitio evades this framework.

svantanaAug 12, 2026
I would be extremely surprised if something as elegant, terse, and useful as the Fourier Transform had been missed by human mathematicians up until now. All expressible theorems are enumerable, after all (if we limit ourselves to a finite alphabet). It seems likely that any new theorems are long, highly complex and esoteric, regardless of human or machine origin.
fn-moteAug 12, 2026
1. The computer is going to struggle to recognize elegance. I’m not sure it’s relevant at this point (but who knows).

2. The statement about proofs is just way wrong. It doesn’t sound like you are familiar enough with them.

This isn’t exactly what you implied, but witness the very short disproof of the Jacobean Conjecture.

Smaug123Aug 12, 2026
Shannon was 1948, one-way crypto in 1978, univalence/HoTT something like 2007; I would be surprised if there weren’t simple new fundamental primitives out there! One problem is that some great advances are from viewing complex objects in a simple way, which take a lot of characters to define in formal logic but which are “simple” in platonic maths-space.
root-parentAug 12, 2026
>> To me, this is more or less what I would think "Artificial General Intelligence" is

So then you need to explain ARC-AGI-3: https://arxiv.org/abs/2603.24621

"Our testing shows humans can solve 100% of the environments, in contrast to frontier AI systems which, as of March 2026, score below 1%."

Back 1996, EQP automatically solved the Robbins conjecture. But nobody concluded EQP was generally intelligent.

https://www.cs.unm.edu/~mccune/papers/robbins/

scronkfinkleAug 12, 2026
> So then you need to explain ARC-AGI-3: https://arxiv.org/abs/2603.24621

I don't, we originally had the turing test which was designed to determine human intelligence by its ability to imitate us with natural dialogue, but we've since defeated that. I stated "to me" because it's my personal opinion on a definition whose goalpost will probably never stop being moved.

> Back 1996, EQP automatically solved the Robbins conjecture. But nobody concluded EQP was generally intelligent.

EQP doesn't have the 3 criteria I outlined, which were different than "solving a math problem"

groundzeros2015Aug 12, 2026
That’s not the criteria outlined in the quote you just used.
seanclaytonAug 12, 2026
Would you trust a bridge built by it with zero human interference? A train? A plane? A skyscraper? What about fight in a war? Unless you can, I wouldn't consider it AGI, because you're actually trusting human intelligence to verify the bridge or train or plane or skyscraper is safe or that the robot is following orders. And even then, you're trusting human-influenced guardrails etc. I would consider it AGI when an AI-created LLM can do all of these things and you trust them with your children's lives. Would you trust an AGI cop to protect your children from a violent criminal? Unless you can, believing what we have as AGI is just an empty opinion with no meaning behind it.
sdenton4Aug 12, 2026
Would you trust a bridge built by a single human? In reality, we have lots of guardrails to ensure that we don't screw up and kill a lot of people by deploying defective bridges (or cars to drive on them). Those guardrails often are written in blood, and still occasionally fail.
seanclaytonAug 12, 2026
Why not have AGI check other AGI? Peer review by fellow humans is what gives humans assurance, to the degree in which review was done by peers of equal or greater intelligence.

AGI checking other AGI should give you that same trust, no? Deepseek says my ChatGPT bridge is stable, you should trust it. Claude says it's stable. The humans say it isn't, but they aren't AGI. You can trust this bridge because it's been vetted by AGI. In my opinion, LLMs cannot be AGI, so for me I would never trust them above any human I would trust. But for those who do believe LLMs can be AGI, they have to demonstrate why we should trust them above any human in these extreme cases. Meaning, if someone says "Well the department of safety (ran by humans) says it's not safe" we have to believe that AGI just knows better than the department of safety. I think this is not possible right now, which is why I don't think we can trust anything built by LLMs where we need the tolerance of risk to human life and safety to approach zero. American AGI soldiers invade the home of Iranian citizens because they have been identified as terrorists. Do you trust the AGI to know if the visual scan they see in this civilian home is a threat to the interests of the United States government and its citizens?

empath75Aug 12, 2026
IME, LLMs are primarily good at grinding through cases, which is why you see them pushing upper and lower bounds and finding counter examples.

I spent a few weeks working on a number theory proof with Claude off and on and it spent hours and hours and hours grinding through one shape of polynomial after another, reporting "progress", and it's true, it proved what I was trying to prove for more and more classes of polynomials, but it was biting off pieces of an infinite tower of classes with no hope of closing it for _all_ polynomials.

That happens to be a good way to find counter-examples, though, and when I posed a slightly different version of my problem, it found a counter example in about 90 minutes.

And in fact, finding the counter example for the related problem allowed Claude to finally prove the thing I wanted to prove to begin with, by lifting the problem to a characteristic where that counter example didn't exist, proving my question there, and then proving that it still was equivalent to my original question.

dataviz1000Aug 12, 2026
If you want to peek inside how a model solves a math problem have a look at some data visualizations I made solving basic multiplication.[0]

I wanted to demonstrate capacity (how well it does a thing) instead of capability (which things it does, like drawing a pelican on a bicycle with SVG or solving a Rubik's Cube). To understand how LLMs solve math, look at the simplest case of multiplication. I deconstructed and classified the thinking token output. It is very important that model training yields thinking token output that structurally follows an observe, orient, decide, act (do the multiplication), and observe again loop.

[0] https://adamsohn.com/reasoning-grid/

kqrAug 12, 2026
You probably know it, but Boyd did not at all suggest speed beats quality. If anything, his realisation was the opposite: the US aircraft had a more open canopy and thus improved quality of observation, and that was what won despite the superior power and turning performance of the Soviet aircraft.
YeGoblynQueenneAug 12, 2026
Disclaimer: I only scanned the article quickly; I might be re-stating something already in the article.

We have just got some very strong evidence about the way in which LLM-based systems solve mathematical problems and this evidence supports what many have already suspected including myself.

Here's what I'm talking about. On 10 August Anthropic released an article [1] claiming that:

An unreleased research version of Claude has improved on a longstanding lower bound for the fraction of zeros of the Riemann zeta function that satisfy the Riemann hypothesis. Drawing on extensive prior research by mathematicians over the past decades, it has increased this bound from 41.6% to 67.2%.

The same article describes the methodology followed by Anthropic's employee, Jarred Sumner, who prompted Claude, as follows:

Jarred Sumner, an Anthropic staff member (and non-mathematician), prompted Claude to “take a real stab” at the hypothesis itself, leaving the mathematical choices from there up to the model. Initially, Claude generated and tried 650 ideas, none of which worked. Jarred prompted Claude to try again, and it spent a day and a half coordinating about 60 Claude subagents, which this time went much deeper: between them, they ran 2,400 shell commands and wrote hundreds of Python scripts.1 The subagents ran thousands of numerical checks against known zeta zeros and refereed one another’s work. Throughout this process, Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.

Jarred got Claude to throw stuff at the wall repeatedly (650 initial "ideas" plus unspecified more by "60 Claude subagents" ... running "2400 shell commands" and "hundreds of Python scripts") and then kept whatever happened to stick. In this case, by happy accident, what stuck was an improved bound of the zeroes of the zeta function etc.

This is how every single mathematical result reported by an AI company has ever been generated. They throw stuff at the wall and take whatever happens to stick.

This approach works. Not only it works, it is, in principle, a universal problem solver. "Millions of monkeys on typewriters" will eventually produce a proof of the Riemann hypothesis; or a disproof of it.

The key point being "eventually". Is this a way to do mathematics research? Can that replace mathematicians?

In AI, this method is well-known as the "generate-and-test" method. It is ancient, basal to AI if I may be so bold. It first appeared to my knowledge in the Logic Theorist, the proof-finding program that Simon and Newell presented in the 1956 Dartmouth convention that named "Artificial Intelligence", to such luminaries of AI and CS as John McCarthy (the real "godfather of AI" who named the field), Marvin Minsky, Claude Shannon and others.

We've had the ability to brute-force all of mathematics "eventually", given "enough" compute for nearing a century now. Why haven't we solved all of mathematics? Are LLMs really so special that they can out-brute force search every previous brute force searcher?

Well, you tell me, HN. I say: no.

___________

[1] https://www.anthropic.com/research/riemann-zeta

yorwbaAug 12, 2026
> We've had the ability to brute-force all of mathematics "eventually"

In a way, yes. You can easily write a program that recursively enumerates all provable theorems in some order. But if you want a proof of a specific theorem, how do you find it in the list? You need to encode the theorem in a formal syntax first, and since mathematics is built on towers of definitions referencing other definitions, that alone is a significant amount of work before you can even write down what you want to prove.

If you want brute force alone, specialized solvers are likely a better choice than LLMs, but what LLMs add to the table is the ability to work with mathematics as it has already been written down. And even though they're bad at brute-forcing, they're still better at it than humans.

An example of a good division of labor is the SAT Attack on Tarski's High School Algebra Problem https://arxiv.org/abs/2608.08421 where they construct a formula with O(n⁴) variables and O(n⁶) clauses and use a SAT solver to show that it is unsatisfiable for n ≤ 11 but satisfiable for n = 12. Then they use an LLM to help them write a Lean proof that the SAT solver input is equivalent to the human-readable description of what they wanted to prove.

YeGoblynQueenneAug 12, 2026
>> In a way, yes. You can easily write a program that recursively enumerates all provable theorems in some order. But if you want a proof of a specific theorem, how do you find it in the list? You need to encode the theorem in a formal syntax first, and since mathematics is built on towers of definitions referencing other definitions, that alone is a significant amount of work before you can even write down what you want to prove.

My current framing of this is that the advantage of LLMs lies in their ability to generate the text of a proof without having derived each of its steps in order, like a theorem prover (automated or not) would have to. There's nothing forcing an LLM to derive conclusions from premises (or indeed making it at all capable to do that).

They don't have to understand what the proof they generate means, or to be able to tell whether it's true. In fact, they can't do either. But that's fine as long as it's possible to check the proof with an external verifier.

So most LLM-based proofs use the LLM as the generator and an external verifier as the tester, either a solver like Lean or a mathematician. That's the best of both worlds as far as generate-and-test goes. A powerful generator tied to a powerful tester.

EDIT: yeah, like this:

>> An example of a good division of labor is the SAT Attack on Tarski's High School Algebra Problem https://arxiv.org/abs/2608.08421 where they construct a formula with O(n⁴) variables and O(n⁶) clauses and use a SAT solver to show that it is unsatisfiable for n ≤ 11 but satisfiable for n = 12. Then they use an LLM to help them write a Lean proof that the SAT solver input is equivalent to the human-readable description of what they wanted to prove.

I'm not disergarding the fact that LLMs don't generate text completely at random. They generate likely text. I suspect that can make it more likely to generate the text of some correct proofs. But I have no idea how likely that "more likely" is or what proofs are those.

ozgungAug 12, 2026
This is not "brute-force" though. It's an iterative search algorithm. You learn things at each iteration. You also don't search blindly. You use "something" (heuristics, experience, intuition) to come up with "ideas" at each iteration. You don't try 650 random programs. You try 650 different ideas each learning from the results of previous trials.

Yes this is the "Universal Problem Solving Algorithm". It's actually the same algorithm used by Evolution.

Also this algorithm is vastly different than "monkeys with typewriters". Monkeys don't learn or evolve their writing. There is no memory, no constraints, no learning-curve. At each iteration they freshly sample from a Uniform Distribution. Expected time for a solution is infinitely long.

"The Universal Algorithm" on the other hand is incredibly fast. Humans (designers, researchers) also use the same algorithm but they are much slower to iterate than computers. Instead of trying 650 different ideas at a single run, we have 100s of researchers each try few different ideas independently.

famouswafflesAug 12, 2026
Thank you. I've never seen 'brute force' and 'type writing monkeys' abused so much than these LLM discussions.
lg5689Aug 12, 2026
LLMs are indeed very special compared to past efforts at automated theorem proving. The search tree for proofs is enormous, even short textbook exercises (i.e. a few dozen lines of Lean) were difficult with GOFAI techniques. Adding a few orders of magnitude to your compute budget barely moves the needle, since the search space increases exponentially for every line of the proof.

Now LLMs have produced multi-thousand line Lean proofs. This is impossible by simply "try everything and see what sticks". LLMs are able to target their efforts to only promising proof strategies. Yes it helps that they work at superhuman speed, so they can try thousands of strategies where a human might try a dozen. But their results cannot be explained only by compute increases; they need genuine mathematical insight.

telAug 12, 2026
Increasingly, I've begun to think of LLMs as sources of really interesting random objects: large pieces of "reasonable thinking" conditioned on a task. It's not that these are correct, in general, but instead they're a concentrated form of random search where that "randomness" is very likely to follow plausible, human patterns.

You can toss it at a task with a suitable machine for transforming that raw material into action and it'll rattle through and sample "plausible human behavior" at that endpoint.

There are more clever ways to use it, but a general tool here is to upgrade any sort of stochastic search to use this new form of random sampling. It'll be way more efficient, properly conditioned, because it just won't visit implausible things nearly as often as competing random sources.

najmbajwa123Aug 12, 2026
From my experience, you have to be good at math to trust an LLM to do the math.
jerfAug 12, 2026
There is an old koan in the old hacker literature:

"A novice was trying to fix a broken Lisp machine by turning the power off and on.

"Knight[, one of the principle designers of the Lisp machine], seeing what the student was doing, spoke sternly: 'You cannot fix a machine by just power-cycling it with no understanding of what is going wrong.'

"Knight turned the machine off and on.

"The machine worked."

I feel like AI is manifesting this even more concretely. I don't feel like I'm guiding the AI super intensely as I work on it with software engineering. I'd have a hard time pointing you at where in the prompt my decades of experience are manifesting. But I definitely can have better results, even with a less frontier-level AI, than people who don't know the same amount of stuff.

Terence Tao also released some unedited transcripts of some of his conversations with AI, and many people observed that while many mathematicians may have been able to formulate the initial question, very few people could have given the same feedback to the AI.

Perhaps someday AI will eliminate the need for competence to use it properly. But that day is not today. And to be honest, that tech is probably not LLMs, no matter how large they get. Some other breakthrough will be necessary to truly eliminate the human element. Those psychopathic elites making plans to turn Earth into one of the Spacer worlds from Asimov's works with a small elite population supported entirely with robots take notes... it's not possible yet.

steinwindeAug 12, 2026
For a list of AI accomplishments in mathematics see https://mathoverflow.net/questions/502120/examples-for-the-u... - or a candidate list here: https://aimath.robertj1.com/ . Many have observed an affinity of AI to the search for counterexamples - or examples. Looking at afore lists, something much more sociological crosses my mind: There is a hunt for answering prominent, clearly stated problems. I'm not a mathematician, but is this mostly what progress in mathematics is about? How about stating worthwhile problems in the first place? What about theory building? Am I right saying this is equally important, but none of those utilizing AI for mathematics seem to be interested in such?
user43928Aug 12, 2026
I know nothing about mathematics, but are there not famous mathematicians like Terence Tao who utilize AI and are obviously interested in theory building?
jerfAug 12, 2026
Given coding agent's demonstrated difficulties with concurrent code, even relatively simple concurrent code, it would be interesting to see how they do with temporal logic. I don't know enough to throw AI at the problems in that space but I wonder if they wouldn't crash and burn on it.

(I haven't had the opportunity to throw a current-gen frontier model at a concurrent problem because I haven't had one to try out lately. The best concurrency is no concurrency and the second-best concurrency is the "web request" model where many web requests are nominally running concurrently but they are otherwise fully isolated from each other and not trying to communicate at all. So maybe they're better, but I feel like if they were a lot better somebody would have noted that in a place I'd have seen by now.)

igor_nastAug 12, 2026
They are exceptional at the spending cost math :)
parhamnAug 12, 2026
> If they were, then their big speed advantage over us would mean that there would be much more of a flood of results.

Is this true right now? Just recently Jarred Sumner tweeted [1] that he managed to make some progress on the Riemann hypothesis while on a jog. Managed to get somewhere by encouraging the llm to “keep going” and “believe in yourself”.

This raised a few questions for me. Had no one at Anthropic thought to try this earlier? It's an interesting footnote that a software engineer there pursued this. How many people in the world can actually verify a proof? How many would we need to sit around and do the right incantations to get a proof out of it? How many would we need to verify and give those proofs value and meaning? What happens when there are more proofs than verifiers? How many will be around in 100 years?

I think it just turns out that a lot this stuff is more socially useful than anything else. The 10 proofs drop came and went in the daily news cycle. Perhaps math is already in it's chess like "for fun" period. I am interested in when we find a very high real-world utility breakthrough math/physics, some space where we've already poured our best human resources at it.

[1] https://x.com/jarredsumner/status/2086869681785500011?s=20

bwfan123Aug 12, 2026
From the tweet:

> Still not sure what that means, but some analytic number theorists seem excited

ie, I prompted AI and it put out a giant pile of tokens. I dont know what it means, but I hope someone gets excited. Mathematicians are now the priests and shamans chanting incantations and taking the holy blessings from the AI gods.

parhamnAug 12, 2026
Thats the issue, they're going to get flooded soon with these holy blessings, then what?
HarHarVeryFunnyAug 12, 2026
It seems intuitive that finding a counter-example might be easier than proving a generality, since you're starting from a concrete goal ("build a foo that has properties X, Y & Z") that you can branch out from, identify sub-problems, etc.

Proving a generality seems much more difficult since you don't know what you are trying to build, although I suppose in some cases you can prove it by proving that it's impossible to construct a counter-example.

BeetleBAug 12, 2026
Since no one has mentioned it yet - just want to point out that Timothy Gowers is a Fields medalist.
richard_chaseAug 12, 2026
That means the post was good, even though the blurry math images are illegible. But yes, as a Fields medalist he is an expert on LLMs.
TormentNexusAIAug 12, 2026
The biggest win for AI dev efficiency is cutting down what gets loaded into context. Semantically matching tasks to the top tools helps a lot.
D4HaAug 12, 2026
Correct me if I'm wrong, this is not the right way to ask this question.

LLMs are good at pattern recognition, so its less a type of math that they'll be good at, and more that when you provide documentation or text that can be easily parsed/compared to its training data/reasoning ability, the better answers you get from an LLM.

Also, you need to be knowledgeable at the same thing you are asking the LLM to do, to verify the answer it gives you (at least for the time being).

justanotherjoeAug 12, 2026
How i reason about this is that, there are two types of science works, exploitative and generative. Exploitative science is kinda like what ingen do in jurassic park, finding an use case based on available tensions in the literature.

Funnily enough, 'generative' ai is not good at generative science, of which requires unique human perception that is not purely symbollic manipulation, but requires a form of revelation. That I think is not here yet with ai...

GPersonAug 12, 2026
There’s a well-known essay from the 90s where Timothy Gowers predicted that the creative/intuitive side of mathematics would be entirely taken over by machines before 2100.