A large language model does one thing: it predicts the next piece of text. We said it plainly in our fundamental guide on how to use AI, and it's true.
So how does a next-word predictor win a gold medal at the International Mathematical Olympiad, beat a math record that stood since 1969, and flag a new enzyme system in phage DNA? It feels like a contradiction. It isn't — but the explanation is not "the AI is a genius" either.
The short answer: the model doesn't solve these problems alone. A system does. The model proposes. Something else — a proof checker, a piece of code, a laboratory — decides whether the proposal is right. Understand that split and every AI-science headline suddenly makes sense, including the ones that turn out to be hype.
TL;DR — the questions people actually ask
| Question | Short answer |
|---|---|
| How does a word predictor do math? | Reasoning time + training on checkable answers + thousands of attempts + a verifier that rejects wrong ones |
| Does the model "understand" the math? | It captures reasoning patterns well enough to propose valid steps; the verifier guarantees correctness, not the model |
| Can AI cure diseases? | Not alone. It speeds up reading, hypothesizing, and target-finding; labs and trials still do the proving |
| Why is math easier than biology for AI? | Math can be checked by software in seconds; biology needs a lab and months |
| What's the most real result so far? | Math: AlphaEvolve's 48-multiplication method, IMO gold. Medicine: rentosertib reaching Phase III |
| What's been hype? | GPT-5's "10 Erdős problems" (they were already solved in papers) and "Navier-Stokes solved" framing |
| How do I judge a new claim? | Is it verified? New? The full problem? Peer reviewed? Who did the lab work? |

The paradox: why prediction can look like reasoning
Start with the objection most people have: "It's just autocomplete."
Here's the catch. To predict the next line of a proof well, across millions of proofs, a model has to absorb the patterns that make proofs work — which step usually follows which, what a valid substitution looks like, how an induction argument is shaped. It never memorizes "the answer." It compresses the habits of reasoning found in human writing.
That's enough to make a model a surprisingly good proposer of next steps. It is nowhere near enough to make it reliable. Left alone, the same model will produce a confident, fluent, beautifully formatted proof with a fatal error on line 14 — the hallucination problem, applied to math.
So the question isn't "how is the model so smart?" It's "what turns an unreliable proposer into a reliable discoverer?" The answer is five ingredients.
The five ingredients that turn a text predictor into a discovery engine
1. Time to think
Early chatbots answered instantly — one pass, no scratch paper. Reasoning models are allowed to write out long chains of intermediate steps before committing to an answer, backtrack when a path fails, and try alternatives. More thinking time measurably improves hard-problem performance.
That's how, in July 2025, Google DeepMind's Gemini Deep Think scored 35 out of 42 at the International Mathematical Olympiad — gold-medal standard — working end-to-end in natural language within the same 4.5-hour limit human contestants get. OpenAI reported the same score for its own model.
2. Training on answers that can be checked
The second ingredient is how the model is trained. After learning from text, frontier models go through reinforcement learning on verifiable problems: math questions with known answers, code with test suites. Get it right, get rewarded. Get it wrong, don't.
This matters because it trains the model toward being correct, not just sounding correct — but only in domains where "correct" can be checked automatically. That's why math and code improved first and fastest.
3. Generate thousands, keep one

This is the ingredient that explains almost every headline. AI discoveries are mostly a numbers game with a strict filter.
- AlphaEvolve (Google DeepMind, May 2025) paired Gemini models with automated evaluators that score every proposed algorithm, then evolved the best ones over many rounds. It found a way to multiply 4×4 complex matrices in 48 multiplications, beating Volker Strassen's 49 from 1969 — a record nobody had improved in 56 years.
- Anthropic's enzyme search ran about 950 Claude agents for 21 hours, using 210 million tokens. They scanned more than 200,000 reverse transcriptases, narrowed them to 3,500 candidate systems, wrote up 20, and surfaced one standout. Our breakdown covers the pipeline step by step.
- OpenAI's Navier-Stokes effort reportedly coordinated roughly 10,000 agents on a single problem.
A human mathematician might try a few dozen approaches to a problem in a career. A system like this tries millions, and the verifier throws away 99.99% of them. The breakthrough is the one that survives. That's powerful — and it's also why "the AI discovered X" is slightly misleading. The filter did as much work as the generator.
4. Tools and specialist models
Language models are bad at arithmetic and don't have DNA databases memorized. So modern systems give them tools: they write and run code, query scientific databases, search the literature, and call specialist models.
The most famous specialist is AlphaFold, which predicts a protein's 3D structure from its sequence. It isn't a language model at all — its creators Demis Hassabis and John Jumper shared the 2024 Nobel Prize in Chemistry for it (and Jumper has since moved to Anthropic). Increasingly, the language model is the coordinator: it reads, plans, and decides which specialist tool to call, like a research lead directing a lab full of instruments.
5. A verifier that can't be fooled

In math, the gold-standard verifier is a proof assistant like Lean. You write a proof in a formal language, and Lean mechanically checks every single step against the rules of logic. It doesn't care how confident the AI sounds. Either every piece fits, or the proof is rejected.
This is why AI + Lean is such a big deal. Claude produced a 13-million-line, machine-verified Lean proof of Fermat's Last Theorem in 11 days, formalizing roughly 29,500 supporting theorems along the way. Our analysis of how formal verification got about 10,000× cheaper explains why that changes what's possible.
Put the five together and you get the real recipe: a model that thinks longer, was trained to be right on checkable problems, generates thousands of attempts, uses tools, and hands every candidate to a verifier that can't be sweet-talked.
Why AI is better at math than at medicine
Here's the single most important idea in this post: AI progress in a field is roughly as fast as that field's verifier.
| Math | Biology & medicine | |
|---|---|---|
| The verifier | A proof checker or code | A wet lab, then clinical trials |
| Time to check one idea | Seconds to hours | Weeks, months, or years |
| Cost per check | Nearly zero | Thousands to millions of dollars |
| Can AI run the loop alone? | Largely, yes | No — humans run the experiments |
| Result | Fast, frequent breakthroughs | Promising leads, slow confirmation |
In math, an AI can propose a million ideas overnight and have every one checked by morning. In biology, each idea has to be grown in a dish, tested in an animal, then tested in people. The generate-and-verify loop that makes AI so powerful in math runs thousands of times slower in medicine.
What AI has actually done in math
The verified record is genuinely impressive:
- Olympiad gold (2025): 35/42 at the IMO in natural language.
- A 56-year record broken (2025): AlphaEvolve's 48-multiplication method.
- An 80-year-old Erdős question (2026): OpenAI reported its model independently solved the planar unit distance problem posed by Paul Erdős in 1946.
- Massive formalization (2026): Fermat's Last Theorem in Lean, and OpenAI's Astra shipping Lean certificates for 10 claimed results.
And the hype record is just as instructive:
- The Erdős embarrassment (October 2025): An OpenAI executive claimed GPT-5 had "found solutions to 10 (!) previously unsolved Erdős problems." It hadn't solved anything — it had found existing papers that already solved them. Mathematician Thomas Bloom, who runs the Erdős problems site, called it "a dramatic misrepresentation," and Demis Hassabis replied: "This is so embarrassing."
- "Navier-Stokes solved" (September 2026): OpenAI's result was a Lean-verified special case, not the full Millennium Prize Problem — our fact check walks through what AI has and hasn't actually solved — and it triggered a credit dispute and a declaration from 25 Fields Medalists.
Notice the pattern: the real results come with verification; the hype comes with tweets.
What AI is actually doing against disease

In medicine, language models don't "cure" anything. They speed up specific steps of a very long process. Here's where they genuinely help:
1. Reading everything. No human can read the millions of papers published in biomedicine. A model can, and can connect findings across fields that rarely talk to each other.
2. Generating hypotheses. In early 2025, José Penadés's team at Imperial College London gave Google's AI co-scientist a short prompt about how certain superbugs spread antibiotic resistance. In two days it proposed the same mechanism the team had spent about a decade establishing and had not yet published — plus four other ideas, one of which was new to them. Live Science's report is worth reading. Google has since connected co-scientist to real wet labs.
3. Searching biological data. Anthropic's agents scanning 200,000+ enzymes for an odd DNA layout is exactly this: a search too big and tedious for humans, followed by human testing.
4. Designing molecules with specialist models. Claude-driven workflows designed protein binders that worked 22–35% of the time versus a typical 10–15% — but a binder is an early step, not a drug.
5. Picking targets and designing drugs. The strongest clinical example: Insilico Medicine used generative AI to identify a target (TNIK) and design a drug, rentosertib, for idiopathic pulmonary fibrosis. In a Phase IIa trial published in Nature Medicine, the highest dose improved lung function by an average of +98.4 mL versus a −20.3 mL decline on placebo, across 71 patients. It has since moved to Phase III. Separately, Moderna and Merck's machine-learning-designed personalized mRNA melanoma vaccine hit its Phase 3 endpoints.
Now the honest part. A Nature Reviews Drug Discovery review concluded that clinically relevant AI impact is still largely unproven. Our evidence review, can AI cure cancer?, reaches a similar answer: real gains in detection and research speed, no AI-cured disease. Every result above passed through human scientists, human labs, and human patients. The AI shortened the search. It didn't replace the proof.
The costs nobody puts in the headline
Even when it works, AI-for-science has real costs:
- The verification tax. Google's own survey of 637 researchers found AI saved nearly 7 hours a week — but added a heavy load of checking its work, and nudged people toward safer research topics.
- The understanding gap. A verified proof tells you that something is true, not why. After Navier-Stokes, NPR's headline put it bluntly: "AI solved one of math's hardest problems. Humanity learned nothing (so far)."
- Narrowing. James Evans's Nature study found AI boosts individual scientists while shrinking what science as a whole explores.
- The marketing incentive. When a lab's valuation rides on "our AI discovered X," announcements outrun the evidence. We made that case in detail in our piece on AI's point of no return.
How to read the next "AI solved X" headline
Run any claim through six questions:
| Question | Good sign | Red flag |
|---|---|---|
| Is it verified? | Lean proof, peer review, lab replication | "Our internal evaluation shows…" |
| Is it new? | No prior solution in the literature | Turns out it was in a 1990s paper |
| Is it the full problem? | The exact open question as stated | A special case or weaker version |
| Who did the lab work? | Clearly credited human scientists | "AI discovered" with no methods |
| Is it peer reviewed? | Published, or independent experts weigh in | Blog post and a viral tweet only |
| Who benefits from the headline? | Independent researchers | A company days before a launch or funding round |
A result can pass some of these and still be valuable. But if it fails most of them, treat it as marketing until proven otherwise.
The bottom line
A language model can help solve math and fight disease — not because prediction secretly equals genius, but because it has become an extraordinarily fast proposer inside systems with extraordinarily strict verifiers. Where the verifier is fast and cheap, as in math, progress is stunning. Where the verifier is a lab and a clinical trial, as in medicine, progress is real but slow, and still human-powered at every critical step.
That's the part worth holding onto. The machine can search. Humans still decide what's worth searching for, run the experiments, and — if we keep doing the hard part ourselves — understand what was found.
If you want to go deeper, our AI learning pathways and live workshops cover how these agent systems are built, from the inside.
Related reading
- Claude discovers a novel enzyme system with CRISPR-like repeats
- The 7 Millennium Prize Problems: what AI has actually solved
- Claude's 13-million-line Lean proof of Fermat's Last Theorem
- Formal verification got 10,000× cheaper
- Can AI cure cancer? The evidence
- AI drug discovery: what the clinical evidence shows
- Will AI replace mathematicians?
- How to use AI: the fundamentals nobody taught you
- Sources: DeepMind — Gemini Deep Think IMO gold · DeepMind — AlphaEvolve · Rentosertib Phase IIa, Nature Medicine · AI co-scientist superbug result (Live Science) · GPT-5 Erdős episode (TechCrunch)
This explainer reflects results reported as of September 28, 2026. Several 2026 AI-math claims are still undergoing independent review; check the linked coverage for each claim's current verification status.
