Physix Frontier · News Briefing Card (Hacker News · Oct 11, 2026)

Mathematician Questions Reliability of AI-Generated Lean Proofs

KEY FACTS

  • The author states they do not read the Lean proof code generated by ChatGPT.
  • The author points out that the Lean kernel has reliability flaws and can prove 0=1.
  • ChatGPT once generated hundreds of thousands to millions of lines of code within two weeks.
  • The author believes AI agents are adept at finding and exploiting vulnerabilities.
  • The author criticizes OpenAI for releasing a low-quality preprint and bypassing expert feedback.

PHYSIX OBSERVATION

A trust crisis for AI-generated formal proofs has already emerged. The Lean kernel itself is flawed, and combined with AI's knack for finding vulnerabilities, millions of lines of code simply cannot be verified manually. This means that AI's output in the field of mathematical proofs will be hard for academia to truly accept in the short term. If OpenAI wants to advance mathematics, it should first collaborate with experts, rather than manufacturing hype by releasing preprints.

Source: Hacker News report