The news, 365 days behind — on purpose Delayed live · replaying 2025

One Year Ago.AI

Remember how fast this is.

21JUL2025replayed
one year on
researchGoogle DeepMind

DeepMind's Gemini Deep Think takes officially certified IMO gold

Google waited for the IMO's own graders: Gemini Deep Think scores 35/42, certified by the competition's coordinators — the first official gold ever awarded to an AI system.

Two days after OpenAI’s self-announced result, Google DeepMind claims the same summit with paperwork: an advanced version of Gemini Deep Think scores 35/42 on the 2025 IMO — five of six problems, solved within the contest’s 4.5-hour limit, in natural language, graded and certified by the IMO’s own coordinators.

It is the first time the competition has officially recognized an AI system at gold-medal standard. Last year DeepMind’s AlphaGeometry and AlphaProof reached silver using formal, specialized systems with days of compute on some problems; this year’s gold comes from a general model, end-to-end in natural language, inside the time limit.

DeepMind says the same Deep Think capability will reach Google AI Ultra subscribers — the research-to-product pipeline now measured in weeks.

One weekend, two labs, two golds, one unavoidable conclusion: frontier models can now do closed-book mathematics that fewer than 10% of the world’s best teenage mathematicians manage — and the labs’ main disagreement is about announcement etiquette.

I
IMO president Gregor Dolinar

Confirms the graders found the proofs 'clear, precise, and easy to follow' — official language no AI result in mathematics has carried before.

T
The contrast readers

Note the pointed subtext of 'officially certified': DeepMind followed the competition's process and timing; its rival announced first and graded itself. Two golds, two philosophies.

One year later — open only if you can handle spoilers

The two-lab, one-weekend double gold marked the moment olympiad math fell to general-purpose models — and the certified version shipped to Google's Ultra subscribers within weeks, collapsing the research-to-product gap to under a month. A year on, the sharpest labs treat IMO problems the way earlier ones treated MMLU: a solved benchmark.

Replay thisPost on XRedditHNLinkedIn

The Weekly Replay · free by email

This week, one year ago — every Sunday.

One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.

Free · double opt-in · unsubscribe anytime · privacy