SOOFI’s German benchmark gains also look inflated by repeated eval exposure
JJitsev · x · 2026-07-26
The author says the German results show the same pattern as the English ones.
- SOOFI trained on all the German rephrased evals, repeated 10 times.
- Nemotron 3 Nano had not seen those evals.
- SOOFI wins on the trained-on German benchmark, but Nemotron leads on HumanEval-DE, again suggesting the comparison is distorted by contamination.
More from Models
- Opus-5 is getting attention for its unusual vocabulary choices — adonis_singh · 2026-07-26
- Google’s Gemma team asks what capabilities people want in the next models — osanseviero · 2026-07-26
- Alibaba answers with Qwen3.8 as Kimi K3 and Chinese models keep closing the gap — emmanuelvivier · 2026-07-26
- Moonshot AI launches open-weight Kimi K3 and claims strong results against top US models — emmanuelvivier · 2026-07-26
- Silicon Valley is split on Chinese open-weight models now rivaling top U.S. systems — emmanuelvivier · 2026-07-26
- Kimi-K3 claims a perfect 6/6 on IMO 2026 Lean 4 proofs — songhan_mit · 2026-07-26