Soofi S 30B-A3B releases a full pretraining report and claims open-model leads in English and German
abursuc · x · 2026-07-21
Soofi S 30B-A3B publishes a full pretraining report, claiming top open-model scores in English and German
The project released the full pretraining tech report and project page for Soofi S 30B-A3B, a Mixture-of-Experts hybrid Mamba model trained on 27 trillion tokens with German deliberately upweighted.
What the report says
- The team says it is the strongest fully open model in its evaluations on both the English and German aggregates, ahead of Olmo 3 32B and Apertus 70B.
- The release emphasizes radical transparency: complete per-source data accounting, all hyperparameters, training and evaluation code, and checkpoints under permissive licenses.
- The model was trained end-to-end on Deutsche Telekom infrastructure.
The thread’s criticism
The quoted reply pushes back hard on the framing, arguing that loudly claiming “sovereignty” and “champion” status creates confusion and misleads the public, especially around German/EU research positioning.
Related event: Soofi S Releases 30B Hybrid Mamba Open-Source Model(2 posts)→
More from Models
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11
- TheZvi Polls: Has Your Coding Model Choice Changed Since Fable 5.1 and Astra? — TheZvi · 2026-09-11