SOOFI Model Accused of Severe Benchmark Contamination in Comparison with Nemotron

The SOOFI model report has been accused of severe evaluation data leakage. Through multiple analyses, author @JJitsev pointed out that SOOFI used unfair, contaminated data when comparing against Nemotron 3 Nano, misleading the public. Once the contaminated benchmarks are removed, SOOFI's actual performance falls short of what the report presents.

Confirmed

Based on @JJitsev's analysis, the following facts can be directly verified from SOOFI's report and data:

- **Training set included eval data**: SOOFI's training set contained multiple evaluation benchmarks later used for comparison. For example, MBPP was included in the training set, while LBPP and HumanEval were not. Table 5 in the report shows that upon removing the trained evaluations, SOOFI's scores dropped significantly.

- **Repeated training on German evals**: SOOFI was trained on all German rewritten evaluations, even repeating them 10 times, whereas the baseline Nemotron 3 Nano was never exposed to this eval data.

- **Partial leaks patched**: In the updated SOOFI-S report, the dev team removed GPQA and dropped the capability index from the previous version. This change caused SOOFI-S's score to decline, roughly falling back to a level comparable with Nemotron 3 Nano.

Unconfirmed

The report labels Nemotron 3 Nano as "open-weights". @JJitsev argues that its data and training stack were already open, questioning whether the SOOFI report exaggerates and misleads regarding its openness and transparency. This subjective assessment remains debated.

Why it matters

The fairness of LLM evaluation is a cornerstone of the research community. If a model has already "seen the answers" during training, its evaluation scores lose their reference value. @JJitsev emphasizes that the SOOFI report attempts to divert attention with marginal patches instead of thoroughly resolving the core issue of eval contamination. This not only misleads the public about its true capabilities but also undermines the fairness of cross-model comparisons.

2026-07-26 ~ 2026-07-26 · 9 related posts

Primary sources