Report: Nemotron fine-tuned to gold level at IMO; IOI run was an unofficial benchmark
NVIDIA and Hugging Face say fine-tuned versions of the open Nemotron model family reached gold-medal level in both of the world’s best-known high school olympiads for math and coding in 2026. But the two results were not documented the same way: the math system was officially entered and graded at the International Mathematical Olympiad, while the coding result on the International Olympiad in Informatics problem set was an unofficial, unsupervised run.
Hugging Face laid out the paired claim in an Oct. 7 blog post titled “One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO,” alongside two NVIDIA technical reports posted to arXiv in September. According to those materials, the math system scored 30 out of 42 points at IMO 2026, above the gold-medal cutoff of 29. The coding system, called Nemotron-3-Ultra-CC, scored 535.4 out of 600 on the IOI 2026 problem set, above the reported gold threshold of 361.12 and above the top human score of 498.27 — but that IOI result was not official.
That distinction is central to the story. The IMO paper says, “The system participated officially in IMO 2026 and scored 30 out of 42 points, above the gold-medal cutoff of 29; all submitted proofs were graded by official IMO graders.” The IOI paper, by contrast, says, “Our system was not an official IOI contestant and the run was not supervised by IOI. Therefore, its score was not included in the official rankings and the evaluation is reported as an unofficial, unsupervised benchmark.”
The underlying claim from NVIDIA and Hugging Face is that one open model family could be pushed to top-tier performance in two very different kinds of reasoning: writing mathematical proofs and solving competitive programming problems under contest constraints. The International Mathematical Olympiad and International Olympiad in Informatics are among the most selective academic competitions for high school students, and a gold-medal score is widely used as shorthand for elite human performance.
These were not base models answering a prompt cold. Hugging Face said, “Starting from Nemotron 3, our teams used supervised fine-tuning (SFT), reinforcement learning (RL), and feedback-driven inference to create systems that reached gold-medal level at both IMO 2026 and IOI 2026.” In plain terms, the systems were specialized after initial training, then further optimized and run with iterative test-time strategies tailored to each contest.
The math setup is especially notable because, according to the IMO paper, it worked entirely in natural language. The system did not use a formal proof assistant, outside tools or internet access during the contest run. That matters because it means the model was judged on the same kind of written mathematical argument that human contestants submit, rather than on machine-checked formal proofs.
Hugging Face and NVIDIA are also framing this as an open release, not just a headline benchmark. The materials being published include model checkpoints, training data, inference code, the submitted IMO proofs and a new 200-problem olympiad benchmark collected on Hugging Face. That makes the result more inspectable than many high-profile AI demonstrations, especially in olympiad-style reasoning.
Still, the caveats are significant. The IOI score should not be described as an official placement, because the paper explicitly says it was neither supervised by IOI nor included in official rankings. The coding system also required enormous compute: the paper says the live IOI 2026 inference deployment used a peak allocation of up to 760 NVIDIA GB300 GPUs. That figure underscores that matching or exceeding top human contest scores can depend not just on model training, but on heavy inference-time resources.
AI companies had already reported major gains on olympiad math in 2025, including DeepMind’s Gemini Deep Think reaching gold-medal standard on IMO-style tasks. What stands out here is a narrower, more concrete combination: one model family posting gold-level results in both math and coding, an officially graded IMO performance, and a public release of the models and artifacts behind the claims. On the evidence NVIDIA and Hugging Face have published, that makes this less a generic model launch than a documented benchmark moment — with one official result, one unofficial one, and an unusually open paper trail for both.
Stocks: NVDA