AI claims math breakthrough: who takes responsibility?
If an AI company says it overturned a mathematical conjecture, should you applaud or take the proof back to check page by page? I'll give my conclusion first: this news is worth attention, but I stand with the mathematicians. The focus is on verification. OpenAI withdrawing sponsorship from Caltech's Math Hackathon puts the credibility issue of AI in mathematics on the table: who announces breakthroughs, who verifies them, and who takes the blame for errors.
Over the past week, tensions between the math community and AI companies exploded. New Fields Medal winner Jacob Tsimerman announced a pivot to AI safety research and joining OpenAI at the International Congress of Mathematicians, while 25 Fields Medal winners jointly expressed dissatisfaction. Earlier, public reports stated that in May, OpenAI used number theory methods to address the Erdős Unit Distance Conjecture. Such news sounds explosive, but the process is more controversial than the result. Mathematicians' complaints are specific: AI companies turn "breakthroughs" into press conferences, while the math community has to recalculate, fix, and take the blame for free.
I actually tried it myself. Recently, I ran some combinatorics and number theory mini-verifications using ChatGPT and DeepSeek-V4-Flash, and wrote a simple enumeration script in Cursor to check outputs in Terminal. Honestly, AI gives ideas quickly, especially good for breaking big problems into finding small examples, guessing patterns, and finally writing scripts to exclude possibilities. It doesn't mind dirty work and can lay out several approaches simultaneously. The problem is obvious: it pretends to understand too well. The language is clean, conclusions are firm, but once a key definition, boundary condition, or step in the proof chain is slightly off, everything afterwards looks correct.
This is also the strongest feeling I got from related discussions on Hacker News these past few days. Many oppose treating "model output" as "mathematical achievement." There's an old rule in math: peer verification decides. Overturning a conjecture requires at least checkable arguments, reproducible reasoning, and challenges from people with different backgrounds. AI's problem lies here: it can generate proofs, but generating a proof and the proof being correct are two different things. Companies can publish blogs, but communities can't rely solely on blogs.
So I view OpenAI's withdrawal from sponsorship more as cutting losses. Joint resistance from new and old Caltech scholars, plus hackathons being a frontier scene for young students to touch AI math, meant continuing to hang the name would turn "sponsorship" into "endorsement." If OpenAI really wants acceptance from the math community, quitting one event is useless; the key is handing over verification rights. Model output shouldn't stop at "I proved it." Ideally, it should provide decomposable assumptions, runnable scripts, and intermediate steps that humans can check line by line. Transparency lowers the heat, but the hassle is real; press conferences aren't as pretty.
This issue connects to the chronic ailment of the entire AI application space. Many AI products sell "results" too early and hide "costs" too deep. Coding, search, and math all hit the same problem. A model gives you code; running doesn't mean safe. AI search gives you a summary; fluency doesn't mean accuracy. Large models give you a proof; neat terminology doesn't mean logical closure. We used to say AI is an assistant; now some companies package it as a referee, which invites trouble.
My judgment is: it depends. At this stage, using AI as a math exploration tool—I recommend it. Using it as an achievement publisher—I don't recommend it. Treating it as mathematical authority—don't touch it. For ordinary readers, upon seeing news like "AI overturns conjecture," don't rush to forward. Find the raw materials, see if anyone reproduced it, see if the community is dissecting details, and see if the company admits uncertainty. For those building AI math products, the action advice is simple: don't just show conclusions; bake verifiability into the workflow. If you can provide runnable code, do it; if you can list assumptions, do it; if confidence is low, label it as low.
The math community's anger this time has conservative elements but also reminds all AI companies: breakthroughs must pass verification to become credit; speed counts only if it passes peer checks. What OpenAI needs to do next is hand over verification rights, making model output checkable line by line by humans.
Physix Frontier