Researcher Flags Data Contamination in SOOFI Model Evaluations
JJitsev · x · 2026-07-19
Researcher **@JJitsev** points out that the SOOFI model, which recently claimed superiority in German evaluations, suffers from severe evaluation data contamination. SOOFI was exposed to most of the evaluation sets during training (including the test set for GPQA), whereas the baseline original Nemotron-3-Nano was not. Claiming "frontier-level" status by comparing this "cheating" method against a baseline model is entirely invalid.
Related event: SOOFI Criticized Over Benchmark Leakage and “Sovereignty” Framing(11 posts)→
More from Models
- Kimi K3 and Fable 5 show nearly identical failure patterns on a software benchmark — FinanceYF5 · 2026-07-21
- Kimi K3 hits 89.4% peak on software tasks while Fable 5 is slightly steadier — FinanceYF5 · 2026-07-21
- Kimi K3 leads on Go, but Fable 5 wins Python, JavaScript, TypeScript and Rust — FinanceYF5 · 2026-07-21
- Kimi K3 costs $4.65 per run and delivers 2.8× more work per dollar than Fable 5 — FinanceYF5 · 2026-07-21
- Kimi K3 reaches 89.4% pass@4 and tops the benchmark over GPT-5.6 Sol — FinanceYF5 · 2026-07-21
- Kimi K3 and Fable 5 now look much closer than the old open-vs-closed gap — FinanceYF5 · 2026-07-21