SOOFI drops GPQA but the Nemotron comparison is still not clean
JJitsev · x · 2026-07-26
The author says SOOFI fixed only one of the leakage issues and the underlying comparison is still invalid.
- SOOFI-S drops GPQA, one of the eval sets it had used in training.
- That also removes the “capability index” used in the previous report.
- Even so, the author argues the remaining comparison is still unfair and now puts SOOFI-S roughly back on par with Nemotron 3 Nano.
More from Models
- Opus-5 is getting attention for its unusual vocabulary choices — adonis_singh · 2026-07-26
- Google’s Gemma team asks what capabilities people want in the next models — osanseviero · 2026-07-26
- Alibaba answers with Qwen3.8 as Kimi K3 and Chinese models keep closing the gap — emmanuelvivier · 2026-07-26
- Moonshot AI launches open-weight Kimi K3 and claims strong results against top US models — emmanuelvivier · 2026-07-26
- Silicon Valley is split on Chinese open-weight models now rivaling top U.S. systems — emmanuelvivier · 2026-07-26
- Kimi-K3 claims a perfect 6/6 on IMO 2026 Lean 4 proofs — songhan_mit · 2026-07-26