SOOFI Model Accused of Severe Benchmark Contamination in Comparison with Nemotron
The SOOFI model report has been accused of severe evaluation data leakage. Through multiple analyses, author @JJitsev pointed out that SOOFI used unfair, contaminated data when comparing against Nemotron 3 Nano, misleading the public. Once the contaminated benchmarks are removed, SOOFI's actual performance falls short of what the report presents.
Confirmed
Based on @JJitsev's analysis, the following facts can be directly verified from SOOFI's report and data:
- **Training set included eval data**: SOOFI's training set contained multiple evaluation benchmarks later used for comparison. For example, MBPP was included in the training set, while LBPP and HumanEval were not. Table 5 in the report shows that upon removing the trained evaluations, SOOFI's scores dropped significantly.
- **Repeated training on German evals**: SOOFI was trained on all German rewritten evaluations, even repeating them 10 times, whereas the baseline Nemotron 3 Nano was never exposed to this eval data.
- **Partial leaks patched**: In the updated SOOFI-S report, the dev team removed GPQA and dropped the capability index from the previous version. This change caused SOOFI-S's score to decline, roughly falling back to a level comparable with Nemotron 3 Nano.
Unconfirmed
The report labels Nemotron 3 Nano as "open-weights". @JJitsev argues that its data and training stack were already open, questioning whether the SOOFI report exaggerates and misleads regarding its openness and transparency. This subjective assessment remains debated.
Why it matters
The fairness of LLM evaluation is a cornerstone of the research community. If a model has already "seen the answers" during training, its evaluation scores lose their reference value. @JJitsev emphasizes that the SOOFI report attempts to divert attention with marginal patches instead of thoroughly resolving the core issue of eval contamination. This not only misleads the public about its true capabilities but also undermines the fairness of cross-model comparisons.
2026-07-26 ~ 2026-07-26 · 9 related posts
Primary sources
- SOOFI’s update still doesn’t fix the benchmark contamination problem — JJitsev · 2026-07-26
- [source] SOOFI drops GPQA but the Nemotron comparison is still not clean — JJitsev · 2026-07-26
- SOOFI appears to have seen several of the benchmarks used against Nemotron — JJitsev · 2026-07-26
- [source] SOOFI’s own table shows weaker scores once you remove training-set benchmarks — JJitsev · 2026-07-26
- [source] SOOFI’s German benchmark gains also look inflated by repeated eval exposure — JJitsev · 2026-07-26
- SOOFI’s English and German scores are both tainted by benchmark leakage — JJitsev · 2026-07-26
- SOOFI may still only be tying Nemotron 3 Nano despite 2T extra tokens — JJitsev · 2026-07-26
- Critics say SOOFI’s updated report overstates openness in a Nemotron 3 Nano clone — JJitsev · 2026-07-26
- SOOFI’s updated report still looks contaminated, and Nemotron 3 Nano may still be ahead — JJitsev · 2026-07-26