SOOFI Accused of Inventing Metrics to Inflate Capabilities
JJitsev · x · 2026-07-19
The author dismisses SOOFI's "frontier-level / champion" claims as baseless. The project allegedly invented a "capability index" that no one else uses to score itself, subsequently claiming to be on par with or even superior to the original Nemotron-3-Nano. He also points out the contradiction with the report's "transparency" narrative. Drawing conclusions based on custom metrics and larger training volumes—without transparently stating these premises—is highly misleading to readers.
Related event: SOOFI Criticized Over Benchmark Leakage and “Sovereignty” Framing(11 posts)→
More from Models
- Kimi K3 and Fable 5 show nearly identical failure patterns on a software benchmark — FinanceYF5 · 2026-07-21
- Kimi K3 hits 89.4% peak on software tasks while Fable 5 is slightly steadier — FinanceYF5 · 2026-07-21
- Kimi K3 leads on Go, but Fable 5 wins Python, JavaScript, TypeScript and Rust — FinanceYF5 · 2026-07-21
- Kimi K3 costs $4.65 per run and delivers 2.8× more work per dollar than Fable 5 — FinanceYF5 · 2026-07-21
- Kimi K3 reaches 89.4% pass@4 and tops the benchmark over GPT-5.6 Sol — FinanceYF5 · 2026-07-21
- Kimi K3 and Fable 5 now look much closer than the old open-vs-closed gap — FinanceYF5 · 2026-07-21