Researcher Flags Data Contamination in SOOFI Model Evaluations
JJitsev · x · 2026-07-19
Researcher @JJitsev points out that the SOOFI model, which recently claimed superiority in German evaluations, suffers from severe evaluation data contamination.
SOOFI was exposed to most of the evaluation sets during training (including the test set for GPQA), whereas the baseline original Nemotron-3-Nano was not. Claiming "frontier-level" status by comparing this "cheating" method against a baseline model is entirely invalid.
Related event: European Open-Source Model SOOFI Accused of Eval Leakage and Overhype(11 posts)→
More from Models
- AI Sextet offers 6 models free and unlimited for 14 days, including DeepSeek and Qwen — airesearch12 · 2026-09-11
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11