German Evaluations Questioned for Skewing Results
JJitsev · x · 2026-07-16
The author points out that Soofi's evidence for being stronger in German also comes from the same batch of rewritten German evaluations. In this case, even if German is added to the training data, it might naturally hold an advantage, making the "stronger" claim unstable.
Overall conclusion: overhyped promotion should stop, otherwise it will damage trust in German/EU outputs and open-source releases.
Related event: SOOFI benchmark claims challenged over leakage and baseline reporting(13 posts)→
More from Models
- Moonshot pauses Kimi K3 signups five days after launch as GPU demand surges — eyishazyer · 2026-07-21
- AI Diplomacy demo makes agents negotiate, ally, and betray each other — jamdac · 2026-07-21
- Newer models need a different prompting style, and old tricks can make outputs worse — emollick · 2026-07-21
- GLM-5.5 is said to arrive in 4 weeks with open weights — tanay_mehta · 2026-07-21
- Fable 5 is credited with a 3-variable counterexample to the Jacobian conjecture — Various-Affect4841 · 2026-07-21
- Ben’s Bites roundup highlights Kimi K3, Fable 5, Cursor costs and self-driving companies — Ben's Bites · 2026-07-21