AI Just Matched the Humans Who Used to Own This Exam
This year, for the first time, an AI system solved every problem on the International Mathematical Olympiad correctly, matching the small group of teenagers who also got a perfect score.
The Scores That Changed the Story
Huawei and Xiaohongshu (also known as RedNote) each said their AI models solved every problem on this year’s IMO correctly, matching the human contestants who hit 100 percent. The two companies were the first AI labs to report a perfect score, though more may follow in the coming days.
Xiaohongshu said a language model had never before achieved a perfect score under the Olympiad’s official judging process. The company’s model, “dots-note-3.0,” was entering the IMO for the first time this year. Huawei’s system, called “Celia,” was described as showing strong problem solving across several branches of mathematics.
How the Test Actually Worked
The AI models did not get an advantage over the human contestants. Companies received the IMO problems only after the students had already taken the exam, then had a set window to submit answers. Xiaohongshu said no human input was allowed once testing began, and the AI’s solutions went to IMO organizers for the same grading the students received.
This year, 666 students competed in Shanghai. Only seven of them walked away with a perfect score, according to the official IMO scoreboard. All IMO competitors must be under 20.
The Gap Closed Fast
A year ago, models from Google and OpenAI reached gold medal level for the first time, a real jump, but they still fell short of the five human contestants who scored 100 percent. Go back one more year, to 2024, and Google’s model was solving four out of six problems over two to three days.
A Second Check, From Outside the Companies
Deedy Das, a partner at Menlo Ventures, ran his own version of the test. He gave this year’s IMO problems to four different AI models, including systems from OpenAI, Anthropic, the startup Axiom Math, and Moonshot AI’s “Kimi K3.” All four scored a perfect 42 out of 42, he said on LinkedIn.
For anyone building a career around these tools, that is worth sitting with. The tasks that felt safely human just got a little smaller.