German Data and Evaluations Skew Comparisons
JJitsev · x · 2026-07-16
This reply corrects a statement in a German model comparison: it should be "better than Nemotron" rather than "versus Nemotron". The author argues that Soofi inherently holds a significant advantage over Nemotron due to having more German training data.
Furthermore, if "rewritten German evaluations" are included in the comparison, using them to prove Soofi's strength over Nemotron becomes invalid. The evaluations themselves are already influenced by training data or language processing methods, distorting the comparative conclusions.
Related event: SOOFI benchmark claims challenged over leakage and baseline reporting(13 posts)→
More from Models
- Unreleased 'GPT 6 Sol' model spotted in OpenAI's API — SteveEricJordan · 2026-09-11
- DeepSeek update keeps cache hits mid-conversation, cuts costs 36.6% — teortaxesTex · 2026-09-11
- AI Sextet offers 6 models free and unlimited for 14 days, including DeepSeek and Qwen — airesearch12 · 2026-09-11
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11