SOOFI Accused of Benchmark Leakage
JJitsev · x · 2026-07-19
SOOFI is accused of eval leakage in its benchmark evaluations. The report allegedly used rewritten evaluation sets—and potentially the test sets themselves—that were already exposed during training, severely undermining its claim of being a "frontier-level champion."
The author points out that looking at LBPP, which was not compromised by the training set, paints a clearer picture: Nemotron 3 Nano scored 38.1, compared to Soofi's 31.0, proving the original model is significantly stronger under un-leaked conditions.
Related event: European Open-Source Model SOOFI Accused of Eval Leakage and Overhype(11 posts)→
More from Models
- Unreleased 'GPT 6 Sol' model spotted in OpenAI's API — SteveEricJordan · 2026-09-11
- DeepSeek update keeps cache hits mid-conversation, cuts costs 36.6% — teortaxesTex · 2026-09-11
- AI Sextet offers 6 models free and unlimited for 14 days, including DeepSeek and Qwen — airesearch12 · 2026-09-11
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11