How Nemotron Was Fine-Tuned For IOI And IMO Gold-Level Results
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: How Nemotron Was Fine-Tuned For IOI And IMO Gold-Level Results on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Hugging Face says two systems based on its Nemotron 3 family scored 535.4 out of 600 at IOI 2026 and 30 out of 42 at IMO 2026. IMO graders officially evaluated the proofs; the IOI result came from an unofficial run and was not included in the competition ranking.

Hugging Face says two systems built from its Nemotron 3 model family achieved gold-level scores at the 2026 International Olympiad in Informatics (IOI) and International Mathematical Olympiad (IMO), as detailed in the original analysis. The IMO proofs received official grading, while the IOI score was produced in an unofficial run and did not count toward the competition’s official ranking.

The programming system scored 535.4 out of 600, above the reported gold threshold of 361.12 and the top human score of 498.27. Hugging Face says it ran prospectively under the same time, internet-access and submission constraints as contestants. However, the company describes the evaluation as unsupervised and unofficial; the score was not part of the IOI’s official standings.

For the IMO, the system scored 30 out of 42 points, above the stated gold threshold of 29, with full credit on four of six problems. Hugging Face says official graders assessed the submitted proofs. The system used natural-language proof generation and revision, without a formal prover, external tools or internet access, according to the company.

The two results came from different specialist setups, not a single identical model and procedure. Hugging Face reports that the IOI system paired supervised fine-tuning of Nemotron-3-Ultra-CC with GenCorrect, an iterative process for generating, evaluating and refining solutions. For the IMO, the company combined its general Nemotron 3 Ultra model with supervised fine-tuning and reinforcement-learning checkpoints.

At a glance
reportWhen: Reported for the 2026 competitions; fur…
The developmentHugging Face reported that domain-specialized Nemotron systems reached gold-level scores in programming and mathematics at the 2026 IOI and IMO.
At a glance
reportWhen: Reported after the 2026 competitions
The developmentHugging Face reported that systems fine-tuned from Nemotron 3 scored above the gold thresholds at IOI 2026 and IMO 2026.

How Fine-Tuning Shaped the Scores

The results offer evidence that specialized training and answer-checking at inference time can help a shared model family handle demanding competition tasks in more than one field. The IOI system had to produce code that could succeed on hidden tests; the IMO system had to produce written arguments that graders judged mathematically valid. Those are distinct measures of performance, and the scores should not be treated as interchangeable.

Hugging Face attributes part of the systems’ performance to repeated evaluation and revision. In the IOI workflow, GenCorrect generated candidate programs and used feedback to improve them. In the IMO workflow, models generated possible proofs, scored and critiqued them, then revised selected attempts. This approach matters because it puts computation into the process of finding and checking an answer, rather than relying only on a single response from a model.

The claims also have different evidentiary weight. Official grading supports the reported IMO score, though the account comes from the team behind the system. The IOI result is a strong company-reported benchmark under competition-like constraints, but it is not an official medal or ranking result. Neither score by itself establishes broad performance on everyday programming or mathematical work.

Amazon

AI programming competition tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From IOI Experiments to Two Fields

Hugging Face presents the 2026 work as an extension of its experiments at IOI 2025. For that event, the company reports that a Nemotron-3-Nano-CC model rose from 130 points before post-training to 280 after supervised fine-tuning and 291 after reinforcement learning. GenCorrect then raised the reported score to 468, above the company’s stated 2025 gold threshold of 438.3. An Ultra-CC version scored 502 using the same test-time strategy, according to Hugging Face.

The 2026 projects applied related ideas to separate competition formats. The IOI work drew on 22,000 programming problems and synthetic reasoning traces. The IMO supervised-fine-tuning data included 414,890 quality-filtered examples from 15,818 proof problems; the reinforcement-learning model was trained on 9,597 problems selected near the model’s capability frontier. Hugging Face says the IMO system drew on complementary strengths in the supervised-fine-tuning and reinforcement-learning checkpoints rather than relying on one alone.

The company’s account says its IMO release includes the two training datasets and model checkpoints, as well as Nemotron-IMO-Bench, a 200-problem olympiad-level benchmark. It also refers to a paper on the IMO training and generate-verify-refine system and a NeMo-Skills repository. The supplied report does not give complete details about every release item.

“Success at both points to something broader.”

— Hugging Face

Amazon

natural language proof generation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of the Competition Evidence

The IOI score was not officially ranked, and the supplied account does not describe an independent audit or replication of the run. Hugging Face says the system operated under the competition’s time, internet and submission constraints, but does not provide enough information here to assess how the run was monitored or how its result was independently checked.

The official grading of the IMO proofs gives that score a different status, but it does not settle how well the system would perform on other proof styles or problems outside this competition. The report also does not establish how either system would generalize to broader programming or mathematics tasks. Results may depend on the selected data, compute budgets and evaluation setup.

Some release information remains incomplete in the supplied account, including details associated with a repository. It also does not provide a timetable for further releases or an independent evaluation of the IOI system.

Amazon

machine learning fine-tuning kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Release Materials and Independent Tests

Researchers can use the IMO checkpoints, datasets and 200-problem benchmark described by Hugging Face to inspect the approach and test it on additional olympiad-level problems. The company also points to a paper and NeMo-Skills materials that describe parts of the training and proof-generation workflow.

The next informative steps would be fuller publication of the methods and release details, independent testing of the systems, and external scrutiny of the IOI evaluation procedure. Until then, the clearest distinction remains: the IMO result was officially graded, while the IOI score was a company-reported, unofficial benchmark run.

Amazon

AI model evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What scores did the Nemotron systems report?

Hugging Face reports 535.4 out of 600 at IOI 2026 and 30 out of 42 at IMO 2026. The two scores came from different specialist systems and evaluation procedures.

Was the IOI result an official medal?

No. The company says the IOI score came from an unofficial run and was not included in the official ranking. It should be described as a reported benchmark score, not an official medal.

How was the IMO result evaluated?

Hugging Face says official IMO graders awarded the submitted proofs 30 points. The company reports full credit on four of the six problems.

What methods did the systems use?

The IOI system used supervised fine-tuning and GenCorrect, which iteratively generates, evaluates and refines solutions. The IMO system combined the general Nemotron 3 Ultra model with supervised-fine-tuning and reinforcement-learning checkpoints, using a generate, critique and revise process.

Can the results be taken as proof of broad ability?

No. The scores show performance on two specific competitions, as reported by Hugging Face. The IOI result was unofficial, and the account does not establish how well either system generalizes to other competitions or real-world tasks.

Primary source: Hugging Face · via ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Experts Warn AI Could Pose A 10%+ Risk Of Human Extinction Following Researcher Departure

A researcher from Anthropic warns there is a more than 10% chance AI could cause human extinction, raising urgent safety concerns.

Understanding The Odin Programming Language

A comprehensive overview of Odin, a new systems programming language gaining attention for its performance and simplicity.

Best Low-Noise PC Cases for Airflow and Sound Dampening

Discover top PC cases balancing airflow and sound dampening for high-performance workloads, with expert insights on choosing the right case for your needs.

NYT Connections today – my hints and answers for June 30 (#1115)

Detailed hints and solutions for NYT Connections puzzle #1115 on June 30, providing insights into the game’s latest update.