A perfect 42/42 at the IMO with no tools: the scaffolding is gone, the reasoning is inside the model

🕒 Published on Zendoric: July 26, 2026 · 00:23
The model scored a flawless 42 out of 42 at the 2026 International Mathematical Olympiad — gold-medal territory — with no external tools and no agent scaffolding. The absence of tools is the headline, not the score.
The fact, per the note: the model achieved a perfect score of 42/42 at IMO 2026, a gold-medal-level result, without external tools or agent scaffolding.
The IMO is not a benchmark that can be memorised. Problems are new each year, proofs are graded by humans, and partial credit is unforgiving — six problems, seven points each, 42 the maximum. A clean sweep means every proof was judged complete and correct, which is rarer than gold itself; most human medallists drop points somewhere.
The qualifier "no tools, no agents" is what makes this different from a leaderboard entry. Systems can be made to look far smarter than they are by wrapping them in search, code execution, verifiers and multi-pass sampling — capability borrowed from the harness rather than resident in the model. Stripping that away turns the result into a statement about the model itself: the multi-step, adversarial reasoning an olympiad demands is now internalised.
Our read: this is one of the cleanest capability signals we have seen, precisely because it is so hard to game — and we would still want the grading protocol and problem-exposure controls published before treating it as settled. The caveat is scope. Olympiad problems are hard but bounded: a correct answer exists, and the question is stated precisely. Research mathematics, and the messy scientific work that actually cures diseases, is mostly the opposite — figuring out which question is worth asking. Systems that can hold a rigorous chain of reasoning together unaided are a necessary ingredient for that, not the finished dish. Short term, expect this to be over-claimed in marketing. Long term, it is exactly the kind of progress that makes accelerated science plausible rather than aspirational.
🔗 Related on Zendoric
- A perfect 42/42 at the IMO: reasoning models now clear the hardest school math without tools · 2026-07-27
- Claude Science: Anthropic turns its model into a lab for researchers and takes direct aim at AI's most ambitious promise · 2026-07-01
- GPT-5.6 clears federal review and ships with 'ChatGPT Work' — the agent race is now about distribution · 2026-07-10


