one year on
OpenAI: experimental model scores IMO gold
An unreleased reasoning model solves 5 of 6 International Mathematical Olympiad problems under contest conditions, writing natural-language proofs graded by former medalists — a milestone many forecasters had penciled in for years later.
OpenAI researcher Alexander Wei announces that an experimental reasoning model has achieved a gold-medal score on the 2025 International Mathematical Olympiad: 35 out of 42 points, solving five of six problems.
The conditions matter: two 4.5-hour sessions, no tools, no internet, proofs written in natural language and graded by former IMO medalists. This is not a specialized theorem-prover with a formal-verification crutch — the claim is that general-purpose reasoning methods got here.
Sam Altman calls it a significant marker of how fast the field is moving, while noting the model is a research artifact: GPT-5 arrives soon, but “we don’t plan to release a model with this level of math capability for many months.”
The milestone instantly generates two arguments. One about timing and manners — announcing during the IMO’s own weekend, with self-graded results, while other labs waited for official certification. And one about calibration: as recently as last year, expert forecasts put this achievement years away. The proofs are public. The forecasts were wrong.
The record
Stresses what's new: no tools, no formal verifiers, no internet — just 4.5 hours, natural-language proofs, and general-purpose reinforcement learning.
Says the model won't ship for many months — GPT-5 is coming soon, but this level of math capability stays in the lab.
Mathematicians and rival labs object to the announcement's timing around the IMO's own closing ceremony, and to self-graded results — a controversy that says as much about the field's trust levels as its math.
One year later — open only if you can handle spoilers
DeepMind's officially-certified gold two days later validated the result wasn't a one-off — and by the following spring, olympiad-level reasoning had shipped in consumer models. The forecasting community's 'AI gets IMO gold' median, which sat at 2027+ as late as 2024, became a case study in systematically underestimating the pace.
The Weekly Replay · free by email
This week, one year ago — every Sunday.
One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.
Free · double opt-in · unsubscribe anytime · privacy