AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Nemotron Fine-Tuning For IOI And IMO: Two Paths To Gold-Level Results on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

Hugging Face says two specialized systems built from its Nemotron 3 model family scored 535.4 out of 600 at IOI and 30 out of 42 at IMO in 2026. The IMO proofs received official grading; the IOI score came from an unofficial run and did not count toward the competition ranking.

Hugging Face says two specialized systems built from its Nemotron 3 model family reached the stated gold thresholds at the 2026 International Olympiad in Informatics and International Mathematical Olympiad, as described in the original analysis. The company reported an unofficial 535.4 out of 600 at IOI and an officially graded 30 out of 42 at IMO, but only the IMO result was assessed by the competition’s official graders.

For IOI, Hugging Face fine-tuned a competition-specific version of Nemotron-3-Ultra-CC using supervised fine-tuning, then paired it with GenCorrect, an iterative system for producing, evaluating and refining code solutions. The company says the system ran prospectively under the competition’s time, internet-access and submission constraints. Its reported score was above the 361.12 gold threshold and the top human score of 498.27, but it was not part of the official ranking.

For IMO, the team combined the general Nemotron 3 Ultra model with supervised fine-tuning and reinforcement-learning checkpoints. The system generated candidate proofs, scored and critiqued them, and revised selected attempts. Hugging Face says official graders awarded the submissions 30 of 42 points, above the stated 29-point gold threshold, with full credit on four of six problems. The company says the system used no formal prover, external tools or internet access.

The projects relied on different training materials. The IOI work used 22,000 programming problems and synthetic reasoning traces. The IMO supervised-fine-tuning set contained 414,890 quality-filtered examples from 15,818 proof problems, while the reinforcement-learning model was trained on 9,597 problems chosen near the model’s capability frontier. These figures and scores are reported by Hugging Face; the supplied account does not provide independent verification of the IOI run.

At a glance
reportWhen: Reported for the 2026 competitions; the…
The developmentHugging Face has reported gold-threshold scores by specialized Nemotron systems in programming and mathematical proof tasks at the 2026 IOI and IMO.
At a glance
reportWhen: Reported after the 2026 competitions
The developmentHugging Face reported that systems fine-tuned from Nemotron 3 scored above the gold thresholds at IOI 2026 and IMO 2026.

Two Different Tests of Specialist Models

The report offers evidence that a shared model family can be adapted to distinct, demanding tasks: producing code that passes hidden tests and writing proofs that human graders judge. The approach combines domain-specific fine-tuning with inference-time processes that check and improve candidate answers, rather than relying on a single model response.

The results carry different evidentiary weight. The IMO score was assessed by official graders, while the IOI result is a company-reported benchmark run outside the official ranking. Even together, the scores do not establish that the systems will perform similarly across other competitions, unfamiliar proof styles or real-world programming work. They show performance on these reported tasks, not general capability across either field.

Amazon

AI coding and proof assistant tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From IOI Experiments to IMO Proofs

Hugging Face describes the 2026 work as an extension of earlier IOI experiments. For IOI 2025, the company reported that a Nemotron-3-Nano-CC model rose from 130 points before post-training to 280 after supervised fine-tuning and 291 after reinforcement learning. GenCorrect raised the reported score to 468, above that year’s stated 438.3 gold threshold; an Ultra-CC version scored 502 with the same test-time strategy.

The 2026 projects applied related ideas to two different evaluation formats. IOI problems require executable programs and performance on hidden tests. IMO problems require written mathematical arguments that graders can assess for correctness. Hugging Face says the IMO team found complementary strengths in supervised-fine-tuning and reinforcement-learning checkpoints and combined them with the general model. That design differs from using one specialist checkpoint alone.

“Success at both points to something broader.”

— Hugging Face

Amazon

machine learning fine-tuning kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Verification and Generalization Questions

The IOI result was unofficial, and the supplied report does not explain how the run was audited or whether an outside group has replicated it. It also does not establish whether the reported competition-like constraints were independently verified. The IMO score has official grader assessment, but further details on how the system’s full process was monitored are not provided in the source material.

It remains unclear how well either system would perform on other contests, new problem sets or practical workloads. The report also does not provide a common evaluation that would make the IOI and IMO scores directly comparable. Hugging Face describes several materials as available, but the supplied account’s closing repository details are incomplete.

Amazon

programming problem datasets

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Checkpoints and Benchmarks for Review

Hugging Face says its Nemotron Labs IMO 2026 collection includes supervised-fine-tuning and reinforcement-learning checkpoints, both training datasets, and Nemotron-IMO-Bench, a benchmark of 200 olympiad-level problems. The company also points to a paper describing the IMO training and generate-verify-refine system, alongside a NeMo-Skills repository.

Those releases could allow researchers to inspect the methods and evaluate the models on additional problems. The supplied source gives no timetable for further releases and does not describe an independent review of the IOI run. Reproduction by outside teams and scrutiny of the IOI evaluation would help clarify how much the reported results depend on the training data, compute and test setup.

Amazon

AI model training and evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Did a Nemotron system officially win an IOI medal?

No. Hugging Face reported a 535.4 out of 600 IOI score, but the run was unofficial and did not enter the competition’s official ranking.

Was the IMO result officially graded?

Yes. Hugging Face says official IMO graders awarded the submitted proofs 30 out of 42 points, including full credit on four of six problems. The company says the stated gold threshold was 29 points.

How were the systems improved?

The projects used supervised fine-tuning and, for the IMO system, reinforcement-learning checkpoints. They also used iterative procedures to generate, evaluate and revise candidate answers.

Can the two scores be compared directly?

No. IOI evaluates programming submissions against test cases, while IMO evaluates written proofs. Their scores measure different tasks, and the IOI result was unofficial while the IMO proofs received official grading.

What materials did Hugging Face say it released?

The company says its IMO collection includes SFT and RL checkpoints, training datasets and a 200-problem benchmark, along with a paper and a NeMo-Skills repository. The supplied source does not specify a release timetable for additional materials.

Primary source: Hugging Face · via ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Scholarship application organizer for school counselors

A new scholarship application organizer for high school counselors is being tested to streamline tracking requirements, deadlines, and student progress.

Transforming U.S. Schools With AI: The Increasing Adoption Of ChatGPT For Teachers

OpenAI is broadening access to ChatGPT for Teachers across more U.S. districts, aiming to support educators with AI tools for lesson planning and grading.

LearnVector – Andrew Ng’s AI Company Building One‑to‑one Learning Experiences

Andrew Ng’s company LearnVector debuts a new platform focused on one-on-one AI-powered learning experiences, aiming to transform personalized education.

The Role Of Attention-Burden Metrics In K-12 Edtech Decision-Making

New focus on cumulative attention-burden scores aims to improve school software procurement by measuring total student attention load across apps.