German Evaluations Questioned for Skewing Results
JJitsev · x · 2026-07-16
The author points out that Soofi's evidence for being stronger in German also comes from the same batch of rewritten German evaluations. In this case, even if German is added to the training data, it might naturally hold an advantage, making the "stronger" claim unstable.
Overall conclusion: overhyped promotion should stop, otherwise it will damage trust in German/EU outputs and open-source releases.
Related event: SOOFI benchmark claims challenged over leakage and baseline reporting(13 posts)→
More from Models
- AI Sextet offers 6 models free and unlimited for 14 days, including DeepSeek and Qwen — airesearch12 · 2026-09-11
- Anthropic publishes its most detailed threat report, including an AI-designed drone swarm case — soumitrashukla9 · 2026-09-11
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11