Fermat's Last Theorem Formalized in 11 Days: How AI Upended Mathematical Peer Review

Fermat's Last Theorem Formalized in 11 Days: How AI Upended Mathematical Peer Review

AIMathematicsAutomation

Sources:Anthropic Research + HN

On September 4, 2026, Anthropic shared an announcement that sent shockwaves through academia: a fleet of AI agents had completed a full machine formalization of Fermat’s Last Theorem in just 11 days. For comparison, a team of leading mathematicians at Imperial College London had secured a £1 million grant and scheduled five full years for that exact undertaking.

The 350-year-old mathematical enigma had already been solved by British mathematician Andrew Wiles in 1995. This dramatic clash—11 days versus five years—does not push the boundaries of mathematics itself.

What it truly dismantles is the centuries-old reliance on human experts to establish epistemic trust. A landmark proof has transitioned from “years of peer review by fallible human mathematicians” to “a machine verifying logical consistency in minutes.”

A 350-Year Mystery Exposes the Limits of Human Review

The history of proving Fermat’s Last Theorem is essentially a record of human review systems repeatedly breaking down. In 1908, the Wolfskehl Prize offered 100,000 gold Marks in Germany for a valid proof. Spurred by the reward, the review committee received 621 purported proofs in the first year alone—every single one of which was found to contain fatal errors.

Human cognition has inherent blind spots when dealing with massive, interconnected logical chains. In June 1993, Andrew Wiles delivered three lectures at the Isaac Newton Institute in Cambridge, announcing his proof to the world. Yet two months later, peer reviewers conducting routine scrutiny posed a sharp question that exposed a critical gap in the proof’s structural machinery.

Wiles spent an entire grueling year repairing the flaw, at times on the verge of abandoning the effort, before finally publishing the definitive 129-page paper in May 1995. Relying on a handful of elite mathematical minds sequestered in a room to examine over a hundred pages of dense manuscript was already pushing human cognitive capacity to its physical limit.

Anthropic Official Announcement Header Figure: Anthropic official announcement header. Source: Anthropic

13 Million Lines of Machine Code in 11 Days

How can a machine verify an intricate web of abstract mathematical symbols? The answer lies in “formalization”: translating loose, prose-style human arguments full of intuitive leaps into computer code where every character and type constraint must strictly hold. Historically, this was tedious, painstaking work carried out only by dedicated specialists.

Professor Kevin Buzzard of Imperial College London had spearheaded a community project planned over five years to formalize Fermat’s Last Theorem. His team produced an 86-page “blueprint” document just to map out the translation roadmap.

Anthropic’s internal research model Claude, running on the Prove2Me platform, compressed that five-year timeline into under two weeks. Following the classic 1995 simplified argument, it generated over 13 million lines of code—more than five times the volume of the platform’s entire official mathematical library.

The codebase encompassed formal proofs for 30,300 lemmas and theorems, with 29,500 incorporated into the final proof graph. Dozens of AI agents collaborated in parallel like tireless gears, breaking down vague human intuition into granular, machine-executable logical chains.

$300,000 to Buy Absolute Certainty

The intervention of industrial-scale compute has transformed high human labor costs into a quantifiable compute balance sheet. Buzzard’s five-year plan budgeted £1 million, reliant on the brainpower, salaries, and prolonged coordination overhead of human scientists.

Claude’s formalization run, by contrast, consumed roughly 6 billion output tokens. Based on current public API pricing, that bill amounts to approximately $300,000. Trading $300,000 of compute for the absolute certainty that previously demanded five years of human labor represents an irreversible shift in the cost structure of frontier research.

Machines do not operate without missteps. In early attempts, the agents suffered coordination breakdowns and lost project state. But the system adjusted quickly through automated trial and error; failed attempts directly contributed roughly 7% of the non-boilerplate lines in the final codebase. Anthropic researchers merely provided high-level specifications, leaving the tactical debugging to the agent swarm.

Proof Route DAG on the Prove2Me Platform Figure: Proof route DAG on the Prove2Me platform, showing how Claude decomposed subtheorems toward FLT. Source: Anthropic

Scientific Trust Handed Over to the Compiler

Professor Kevin Buzzard personally verified the multimillion-line codebase, which passed the compiler with zero errors. Automated comparison tools also confirmed that all theorem statements aligned perfectly with existing formal mathematical libraries. This achievement officially completed the 20-year-old “Formalizing 100 Theorems” benchmark challenge.

As Buzzard himself noted, no new mathematics was uncovered, but the achievement provides powerful empirical evidence for automated formalization. When a 129-page mathematical proof can be compiled just like source code in software engineering, academia’s underlying trust network is irrevocably altered.

Historically, the mathematical community accepted a theorem because authoritative figures vouched for it. Today, that centralized authority model has dissolved. As long as code compiles cleanly, mathematical truth requires no human guarantor. The compiler has stepped in as the ultimate line of defense.

Reference Links:

  • Anthropic Research
  • HN Discussion