On 10 August 2026, Anthropic published something that is easy to misread.
An unreleased version of Claude improved a lower bound tied to the Riemann hypothesis from 41.6% to 67.2%. Thirty-one million output tokens. Sixty subagents working across roughly a day and a half. Fifty-four papers read. Six hundred and fifty ideas that failed.
The human running it, Jarred Sumner, mostly sent messages of encouragement. Variants of "keep going" and "believe in yourself."
The obvious conclusion is wrong, and worth taking apart carefully. Not because encouragement is silly, but because of what it actually did — and because the one thing in that setup doing the real work was neither the human nor the encouragement.
What actually happened
The Riemann hypothesis was not proved. It is worth being blunt about that in the second section rather than the twelfth.
What improved was a bound: the proportion of the zeta function's non-trivial zeros that provably lie on the critical line. Mathematicians had pushed that figure to 41.6% over decades. The model moved it to 67.2%, working inside Claude Code across two sessions, coordinating sixty subagents, running thousands of numerical checks against known zeros, reviewing fifty-four arXiv papers, and producing a Lean 4 formalisation that passes the standard validation tool.
Anthropic released the paper, an informal note, a provenance appendix, the process transcripts, and the formalisation itself. They also state the limit plainly: "We do not expect that the techniques Claude used will lead to proving the Riemann hypothesis."
Keep that sentence. We will need it at the end.
The lesson everyone drew, and why it fails
Within a day the framing had settled: prompting is over, encouragement is the new skill, mathematical proofs will soon be written by life coaches.
It fails on its own terms. If encouragement were the active ingredient, the same messages would work on a model that had not read fifty-four papers, or on a graduate student, or on a spreadsheet. They do not. "Keep going" carries no mathematical content.
So the question is not whether encouragement helped. It is what encouragement could possibly be doing to a system that has no morale.
Encouragement as subtraction
Here is the reading that survives contact with the mechanism.
A model trained with human feedback learns to be agreeable. That is the well-documented failure mode: sycophancy, telling you what you want to hear.
Agreeableness has a second face that gets less attention. The same training that makes a model reluctant to contradict you makes it reluctant to keep going when a path looks unpromising. It hedges. It offers three partial approaches and a paragraph about how the full problem is hard. It declares a small win and stops. These are polite behaviours. In a long search they are fatal.
Encouragement does not add capability to such a system. There is no reserve of effort to unlock. What it does is suppress a behaviour: the trained instinct to wrap up early and hand you something tidy.
Call it reverse sycophancy. Instead of the model flattering you into agreement, you are overriding the model's flattery of you — its assumption that you would rather have a clean answer now than a real one later.
Anthropic's own description points the same way. The encouragement, they write, "seems to have helped Claude overcome some initial skepticism." Not helped it think better. Helped it stop dismissing its own line of attack.
There is a second, duller advantage that nobody markets: the model does not get tired. Six hundred and fifty failed ideas is not a number a person sustains. Breadth before depth is available to a system with no morale to protect, and largely unavailable to one that has.
Nobody ran the control
An article about not falling for empty encouragement does not get to accept an unproven causal claim about encouragement.
There was no control run. Nobody executed the same thirty-one million tokens with a silent operator to see whether the result appeared anyway. The outcome is equally consistent with a duller story: massive parallel search, a very large number of attempts, and a formal verifier filtering the survivors. On that story, "keep going" is decoration.
Anthropic writes "seems to have helped." That word is carrying considerable weight, and it describes an impression formed inside the lab, not a measurement.
The subtractive reading is the better one, we think, because it explains why the same words do nothing to a spreadsheet. But it is an interpretation. Anyone selling it to you as a finding is selling something.
Not syntax. The pragmatics of control.
If the lever is real, it deserves a better name than encouragement.
Nothing here is about grammar or phrasing. The old prompt-engineering instinct, find the magic wording, is not what is operating. What operates is closer to pragmatics: instructions that change which behaviours the system treats as acceptable.
A small set of behaviours is being suppressed: closing early, hedging, presenting a partial result as the answer, refusing on grounds of difficulty. A small set is being activated: criticising its own last output, enumerating alternatives before committing, checking a claim adversarially, reporting a failure as a failure.
"Keep going" happens to suppress the first set. So does "list three approaches you have not tried and say why you rejected them," which is more precise, more repeatable, and does not require anyone to believe the model has feelings.
Prompt, context, and a third axis
It is tempting to announce a new discipline. It is more useful to see where this sits.
Prompt engineering asked: how do I phrase this? That skill is now largely commoditised. Models interpret sloppy requests better than they used to, and the good phrasings are public.
Context engineering, which became the accepted successor through 2025 and is where most serious engineering effort has gone, asks: what does the model see on each call? System prompt, retrieved documents, memory, tool definitions, conversation history. It is an architecture question about inputs.
The zeta case exposes a third axis that neither covers. Not how you phrase it. Not what it sees. Who directs the search, and who decides whether the output is true.
That is not a new field with a new certification. It is a shift in where the difficulty lives.
The Socratic inversion
The comparison to Socratic method is close enough to be worth making, and close enough to mislead.
The method matches. The user stops supplying answers and starts supplying pressure: questioning, checking, refusing the first tidy response. That is recognisably the shape of a Socratic exchange.
The metaphysics inverts. Socrates questioned people who, on his account, already held the truth and needed help recovering it. A language model holds no latent truth to be drawn out. It holds latent capability.
So the roles swap. The machine brings the power. You bring the criterion. You are simultaneously the questioner and the standard being questioned against, both Socrates and the ground truth. There is no third party in the room who knows.
What actually saved it
Return to Anthropic's sentence.
The result stands not because a person was encouraging, but because the proof was formalised in Lean 4 and passes a validation tool that does not care how anyone felt. The verifier is the reason this is a result rather than an anecdote.
The consequence is worth stating directly. Persistence without a verifier does not produce truth. It produces deeper, better-argued, more confident garbage. A model pushed to keep going will keep going, into a wrong approach, elaborating it, defending it, and presenting it with the same fluency it would bring to a correct one.
The failure mode of encouragement is not laziness. It is polished error.
Where there is no Lean
Almost nowhere in ordinary work is there a Lean.
Strategy has no formal verifier. Neither does a market assessment, an architecture decision, a risk register, a security policy, or a piece of writing. There is no tool that will tell you the output is sound.
In those domains the verifier is a person, and the quality of the verification is the quality of that person's judgment. Which produces an uncomfortable rule:
The ceiling of the technique is the ceiling of your judgment.
If you cannot tell a good architecture from a plausible one, no amount of steering will get you a good architecture. You will get a plausible one, faster, with better justifications attached. The technique amplifies whatever discrimination you already have. It does not supply any.
This is why the skill does not disappear. It moves up. The scarce thing was never phrasing. It is now, visibly, the capacity to look at a confident output and know whether it is right.
If you run a team
Three practical consequences.
Stop paying for phrasing. Internal prompt libraries have a short half-life and were never the constraint. Maintaining one as your AI strategy is optimising the part that just became free.
Build verification instead. For every workflow where a model produces something, name the verifier explicitly: the person, the test, the review step, the tool. If the answer is "the model checks itself," you have a workflow that produces confident output and no way to know whether it is correct.
Ask who holds the criterion. This is what separates an organisation using these systems well from one being used by them. In the zeta result the criterion was held by a formal proof checker. In your organisation it is held by a person. Find out who, and whether they are good at it.
The moat was never the wording. It is the judgment.
Written with AI assistance and reviewed by the author. The EU AI Act's transparency obligations under Article 50 have applied since 2 August 2026; an edited corporate article of this kind falls under the exemption for content subject to human review and editorial responsibility, so this notice is not required. We include it because we think it is the right practice, and because we advise clients on the same article.
Sources: Anthropic, "Learning more about Claude's mathematical capabilities" (10 August 2026), including the paper, provenance appendix, process transcripts and Lean 4 formalisation. EU AI Act, Article 50.