Anthropic's Claude Generates 13-Million-Line Proof of Fermat's Last Theorem
Anthropic's Claude AI has generated a 13-million-line formal proof of Fermat's Last Theorem in Lean over 11 days. The machine-verifiable codebase has been published on GitHub for global mathematical review.

The artificial intelligence safety and research company Anthropic has announced a major milestone in computer-assisted mathematics. Its AI model, Claude, spent 11 days generating a complete, machine-verified formal proof of Fermat’s Last Theorem. According to the company, this constitutes the longest mathematical proof ever constructed, spanning 13 million lines of code.
To achieve this, Claude wrote the proof in Lean, a specialized programming language and proof assistant designed to formulate mathematical reasoning in a way that computers can verify step-by-step. Instead of relying on human-readable mathematical prose, Lean deconstructs the proof into rigorous logical steps, eliminating human subjectivity and potential oversight.
The Historical Context of Fermat's Enigma
Fermat’s Last Theorem states that no three positive integers $x$, $y$, and $z$ can satisfy the equation $x^n + y^n = z^n$ for any integer value of $n$ greater than 2. French mathematician Pierre de Fermat famously scribbled the conjecture in the margin of a copy of Diophantus' Arithmetica in 1637, claiming he had a "truly marvelous proof" that the margin was too narrow to contain. He died without writing it down, leaving mathematicians to spend 358 years trying to determine if such a proof ever existed.
An accepted proof was finally delivered in 1995 by British mathematician Andrew Wiles. Wiles initially announced his solution during three lectures in June 1993, but a critical flaw was subsequently discovered. He spent nearly a year correcting the error alongside his former student Richard Taylor, publishing the revised 129-page proof in May 1995. Because Wiles' proof relied on advanced 20th-century mathematics unavailable in the 17th century, most historians doubt Fermat actually possessed a valid proof.
Auto-Formalization and the Multi-Agent System
The experiment was led by Tianyi Peng, who develops AI-driven formalization tools with a research team at Columbia University. Dozens of Claude agents worked in parallel, writing definitions, proving minor lemmas, and linking them into larger mathematical structures. Human intervention was kept to a minimum, limited to high-level task prioritization.
Kevin Buzzard, a professor at Imperial College London who leads the Fermat formalization project, reviewed the output and described it as an extraordinary achievement:
"Auto-formalization takes mathematical proofs written for humans and translates them into code where every single step can be verified by a computer. Claude's success on such a complex proof suggests that this technology may soon formalize vast areas of modern mathematics."
Overcoming Technical Hurdles
The process faced initial setbacks. Early on, the AI agents lost track of completed proofs and failed to coordinate effectively. These failed attempts still account for roughly 7% of the lines in the final output. To resolve this, Peng's team deployed a tool called Prove2Me, which maintained a live task list of remaining lemmas, structured files for faster Lean compilation, and saved natural language notes so agents could build upon each other's work.
Ultimately, Claude proved over 30,000 auxiliary theorems, consuming billions of tokens. The task ran on a research model closely resembling Claude Fable 5.1. The final 13-million-line proof is more than five times the size of Mathlib, the standard community library used for formal mathematics.
Anthropic emphasized that this achievement does not represent the discovery of new mathematics, but rather the formalization of Wiles' existing proof. The complete 13-million-line codebase has been uploaded to GitHub, where it is publicly available for verification.





