SOOFI’s English and German scores are both tainted by benchmark leakage
JJitsev · x · 2026-07-26
The author says SOOFI’s English and German evaluation results are badly compromised by contamination.
- He argues SOOFI trained on many of the benchmarks later used for comparison.
- That makes the comparison with Nemotron unfair, because Nemotron had not seen those evals.
- He points to a mixed result: SOOFI can score higher on benchmarks it trained on, while Nemotron leads on unseen ones such as HumanEval-DE.
More from Models
- Kimi-K3 claims a perfect 6/6 on IMO 2026 Lean 4 proofs — songhan_mit · 2026-07-26
- OpenAI’s “5.6 Pro” is said to beat Fable 5 on math-heavy rebuttal work — MParakhin · 2026-07-26
- Claude keeps defending Opus 5 instead of writing a critical review — sethlazar · 2026-07-26
- Unsloth’s Qwen3.6-35B-A3B GGUF is trending on Hugging Face — unsloth · 2026-07-26
- SOOFI may still only be tying Nemotron 3 Nano despite 2T extra tokens — JJitsev · 2026-07-26
- A simple “ask clarifying questions first” prompt gets Claude to surface missing assumptions — Commercial-Most3081 · 2026-07-26