German Data and Evaluations Skew Comparisons
JJitsev · x · 2026-07-16
This reply corrects a statement in a German model comparison: it should be "better than Nemotron" rather than "versus Nemotron". The author argues that Soofi inherently holds a significant advantage over Nemotron due to having more German training data.
Furthermore, if "rewritten German evaluations" are included in the comparison, using them to prove Soofi's strength over Nemotron becomes invalid. The evaluations themselves are already influenced by training data or language processing methods, distorting the comparative conclusions.
Related event: SOOFI benchmark claims challenged over leakage and baseline reporting(13 posts)→
More from Models
- Moonshot pauses Kimi K3 signups five days after launch as GPU demand surges — eyishazyer · 2026-07-21
- AI Diplomacy demo makes agents negotiate, ally, and betray each other — jamdac · 2026-07-21
- Newer models need a different prompting style, and old tricks can make outputs worse — emollick · 2026-07-21
- GLM-5.5 is said to arrive in 4 weeks with open weights — tanay_mehta · 2026-07-21
- Fable 5 is credited with a 3-variable counterexample to the Jacobian conjecture — Various-Affect4841 · 2026-07-21
- Ben’s Bites roundup highlights Kimi K3, Fable 5, Cursor costs and self-driving companies — Ben's Bites · 2026-07-21