TLDR: For the first time in history, artificial intelligence models from Google DeepMind and OpenAI have achieved gold medal status at the prestigious International Mathematical Olympiad (IMO) 2025. Both advanced AI systems successfully solved five out of six exceptionally complex mathematical problems under standard competition conditions, demonstrating a monumental leap in general-purpose AI reasoning, creativity, and logical problem-solving capabilities.
In a landmark achievement that signals a new era for artificial intelligence, Google DeepMind’s advanced Gemini Deep Think model and OpenAI’s experimental reasoning large language model have both secured gold medal-level performance at the 2025 International Mathematical Olympiad (IMO). This marks the first instance of AI systems reaching such a high echelon in the world’s most prestigious mathematics competition for pre-university students.
The IMO, held annually since 1959, brings together the brightest young mathematicians globally, with only about 8% of contestants typically receiving gold medals. The 2025 competition, which took place in Sunshine Coast, Australia, saw 630 students tackle six exceptionally difficult problems spanning algebra, combinatorics, geometry, and number theory. These problems are renowned for demanding not just computational ability, but genuine creativity, insight, and mathematical intuition.
Both Google DeepMind’s and OpenAI’s models operated under the rigorous constraints of human contestants: two 4.5-hour examination sessions, no access to external tools or the internet, and the requirement to produce complete, natural language mathematical proofs. Each model successfully solved five out of the six problems, achieving a score of 35 out of a possible 42 points. Their solutions were independently reviewed and scored by former IMO judges and medalists, who confirmed their gold medal-level performance.
This breakthrough is particularly significant because it demonstrates AI’s ability to engage in sustained, deep thinking—up to 100 minutes per problem—to construct intricate proofs. Unlike prior AI systems narrowly designed for specific mathematical tasks, these models represent the rise of general-purpose reasoning systems. OpenAI researchers, including Alexander Wei, highlighted that this capability level was reached not through narrow, task-specific methodologies, but by breaking new ground in general-purpose reinforcement learning and test-time compute scaling.
Sebastien Bubeck, a former Microsoft VP AI who recently joined OpenAI, reportedly described this achievement as potentially a ‘moon landing moment’ for AI, underscoring the profound implications for what machines are now capable of. The fact that these general-purpose systems, rather than specialized tools, achieved such a feat suggests that AI capabilities are advancing faster than many expert predictions.
While OpenAI announced their results first, Google DeepMind’s Gemini Deep Think also achieved comparable success, highlighting a competitive dynamic in the AI research landscape. OpenAI CEO Sam Altman has indicated that the company does not plan to release an LLM with these advanced math problem-solving capabilities for several months, suggesting further refinement and integration before public deployment.
Also Read:
- AI Breakthrough: Gemini 2.5 Pro Achieves Gold Medal Performance on IMO 2025 Problems
- DeepMind’s AlphaEvolve: An AI System Revolutionizing Algorithmic Discovery and Optimization
This historic performance at the IMO signifies a clear inflection point in AI development, showcasing its capacity for complex, creative, and logical reasoning that mirrors, and in some cases, surpasses human prodigies. It sets the stage for future AI models with significantly enhanced reasoning abilities, promising transformative impacts across various domains.


