Claude calculates zero-point ratio at 67.2%, but this might not be a math story at all
From 41.6% to 67.2%. If you presented this number at a math department seminar, it would be enough for a young scholar to coast on for ten years. But placed in an Anthropic PR release, I prefer to read it as a business signal regarding reasoning chains, rather than the dying breath of the Riemann Hypothesis.
Let's talk about the breakthrough itself. Pushing the lower bound of the proportion of zeros on the critical line of the Riemann ζ function from 41.6% to 67.2% is solid progress in analytic number theory. For decades, this number has been nudged up bit by bit; every nudge was a paper-level engineering feat. Claude completed this autonomously in just a few days. Regardless of the internal reasoning tricks used, the result itself can be verified using traditional mathematical methods. So I don't doubt the authenticity of the number; what I doubt is how it was derived.
Anthropic hasn't disclosed the reasoning process of this research model, only stating "completed autonomously over several days." In the math community, this reveals the core contradiction. The mathematical community recognizes verifiable proof chains, not a model telling you, "I calculated it to be 67.2%." When Leibniz fought Newton over calculus, they were arguing about methodology, not computational results. Now Claude gives a number, but the entire reasoning chain is a black box. This effectively turns mathematics into an oracle. It reminds me of last month's incident where Fable 5 disproved the Jacobian Conjecture. That also involved an "undisclosed model," and to this day, the math world is still debating how to re-verify it. Putting these two events together, Anthropic's playbook is clear: use hard math problems as stress tests for reasoning capabilities, rather than actually trying to solve the Riemann Hypothesis.
Looking at this from three dimensions, I feel it resembles a declaration of commercial positioning more than anything else.
This is a direct rebuttal to the stereotype that "AI only handles language processing." Mathematical reasoning has long been considered the fortress of symbolic logic, with large models frequently questioned as being nothing more than "advanced pattern matching." By pushing the zero proportion forward significantly, regardless of the process, this result itself is the strongest counter-argument to the "pattern matching theory." Anthropic choosing the Riemann Hypothesis over other problems is intentional. The hypothesis has high visibility among both the public and academia, and it comes with the drama of being "unsolved for a century." Combined with Fable 5's disproof of the Jacobian Conjecture, Anthropic is building a narrative line that "AI can discover new mathematics." If this line is accepted by the market, they are no longer selling APIs, but selling "frontier discovery capabilities."
But there's a catch here: ChatGPT is doing something similar. Last month, ChatGPT solved an Erdős prize problem with a single page. The method is almost identical to Claude's: undisclosed model, autonomous completion, stunning results. Both top labs are using hard math problems as billboards for their reasoning capabilities. This makes me suspect whether the underlying LLM capabilities have hit a common bottleneck, leading everyone to turn to "extreme testing" to create differentiation. If so, the real highlight of this Riemann breakthrough isn't the math, but the verifiability of the reasoning chain.
My own judgment is that this track will split into two routes over the next six months. One is "result-oriented," which I call "black box shock": the model provides conclusions, humans verify them, and if verification fails, it goes back to the drawing board. The other is "explainable reasoning": the model breaks down each step of reasoning into auditable intermediate steps that humans can check one by one. Currently, both Anthropic and OpenAI are betting on the first route, but the math community wants the second. Last week, in Project, I tried having Claude solve a numerical solution for a partial differential equation. It gave me the correct numbers, but when asked to write out the derivation steps, it started getting vague. This experience leaves me with reservations about the gold content of "autonomous proofs."
There's another unexplored perspective: copyright. Anthropic recently settled a $1.5 billion lawsuit with authors, indicating their training data included a lot of copyrighted text. Copyright ownership for math papers differs from literary works, but if the reasoning path overlaps with an unpublished paper draft, this controversy will become very messy later on. However, these are future concerns; rushing to conclusions now is pointless.
Back to the Riemann Hypothesis itself. The number 67.2% is indeed beautiful, but if you ask a friend working in analytic number theory, they'll likely tell you that the real difficulty in this direction isn't the lower bound of the zero proportion, but how to translate the advancement of the lower bound into an understanding of the essential structure of the critical line. In other words, Claude took a big step forward, but it stepped onto a road that had already been validated, rather than blazing a new trail. It's like driving from Beijing to Shanghai: Claude massively increased the average speed, but it didn't discover a new highway.
So my advice is: pay attention to the result, but don't rush to treat it as a milestone for "AI solving millennium problems." The math community will spend months verifying this 67.2%. If the proof process can be broken down into a verifiable reasoning chain, then it's a true breakthrough; if it's just some implicit computation inside the model, it's more of an engineering marvel than a mathematical discovery. For those in the AI industry, what's truly worth watching is whether Anthropic will open up the reasoning logs of this model next. If they do, it shows confidence in verifiability; if they don't, then this 67.2% is just a carefully packaged marketing number.
Based on my testing, Claude's capabilities are indeed rapidly approaching a critical point, but it's still some distance away from being a "trustworthy mathematical collaborator." This gap cannot be filled by compute power alone; it requires transparency and auditability.
Physix Frontier