NVIDIA Paper Says AI Scored Above Top Human at IOI 2026 in Unofficial Run
NVIDIA researchers say an artificial intelligence system scored 535.4 out of 600 on the 2026 International Olympiad in Informatics problem set, topping both the contest’s gold-medal cutoff and the highest human score. But the result came in an unofficial benchmark: the system was not an official IOI entrant, was not supervised by the contest, and did not appear in the official rankings.
The claim appears in an arXiv paper, “Post-Training Language Models for Gold-Medal Performance in Coding Competitions,” posted as arXiv:2609.02849 by Aleksander Ficek, Sean Narenthiran, Mehrzad Samadi, Somshubra Majumdar and Boris Ginsburg. The paper’s PDF is marked “© 2026 NVIDIA.” ArXiv shows version 1 was submitted Sept. 2, 2026, and version 2 was revised Sept. 4, 2026.
According to the paper, the competition-specific system Competition Nemotron-3-Ultra-CC scored 535.4 on the IOI 2026 problem set. The authors say that was above the IOI 2026 gold-medal threshold of 361.12 and above the top human contestant’s 498.27.
The paper also stresses a key limitation. “Our system was not an official IOI contestant and the run was not supervised by IOI. Therefore, its score was not included in the official rankings and the evaluation is reported as an unofficial, unsupervised benchmark,” the authors wrote.
That caveat is central to the claim. The International Olympiad in Informatics is one of the world’s best-known programming contests for secondary-school students, and its problems are widely treated as a hard test of algorithmic reasoning. But this was not a formal contest result.
The authors describe the IOI 2026 run as prospective, meaning it was carried out during the competition window before the problems were publicly released. They say the system operated under the same time limits, submission platform and internet-access restrictions as human contestants. Local code execution was allowed, they wrote, with up to 50 submissions per problem and one submission per minute.
The work used two models. Nemotron-3-Nano-CC is described as a 30 billion-parameter mixture-of-experts model with 3 billion active parameters, trained with supervised fine-tuning and reinforcement learning. Nemotron-3-Ultra-CC, the larger system used for the headline result, is described as a 550 billion-parameter model with 55 billion active parameters, used with supervised fine-tuning.
The paper says the researchers curated 22,000 competitive-programming problems and used an iterative test-time method called GenCorrect to generate, evaluate and refine solutions. During the live IOI 2026 inference run, they report a peak allocation of up to 760 NVIDIA GB300 GPUs.
That hardware figure is important context. The paper says the AI was tested under contest-style rules for time, submissions and internet access, but its compute budget was obviously not comparable to a human competitor’s.
Competitive programming has become a prominent benchmark for AI coding systems. Earlier efforts such as AlphaCode showed strong performance, and more recent work from other groups, including OpenAI, reported gold-level results on some contest-style benchmarks. What makes this paper notable is the authors’ narrower claim: a prospective live run on an IOI problem set that, they say, exceeded the top human score, while still being unofficial.
The authors phrase that carefully. “To our knowledge, this is the first AI system to outscore the highest-scoring human contestant on an IOI problem set,” they wrote.
NVIDIA also promoted the result publicly. In a social post, the company said its fine-tuned Nemotron model scored 535.4 out of 600, “as graded by the IOI team — higher than the top-scoring human participant.” The paper itself, however, says the run was not supervised by IOI and was reported as an unofficial benchmark.
The researchers said they plan to release the competition checkpoint and inference and evaluation recipes through NeMo-Skills, NVIDIA’s toolkit for model post-training and evaluation.
Stocks: NVDA