SOOFI’s own table shows weaker scores once you remove training-set benchmarks
JJitsev · x · 2026-07-26
The author uses the report’s own Table 5 to show how training on eval sets shifts the comparison.
- MBPP was included in SOOFI’s training, while LBPP and HumanEval were not.
- SOOFI scores lower on the evals it did not train on, while Nemotron 3 Nano comes out ahead there.
- The thread also lists the evaluation sets SOOFI saw versus the datasets used to train Nemotron 3 Nano, arguing the comparison is not clean.
More from Models
- Opus-5 is getting attention for its unusual vocabulary choices — adonis_singh · 2026-07-26
- Google’s Gemma team asks what capabilities people want in the next models — osanseviero · 2026-07-26
- Alibaba answers with Qwen3.8 as Kimi K3 and Chinese models keep closing the gap — emmanuelvivier · 2026-07-26
- Moonshot AI launches open-weight Kimi K3 and claims strong results against top US models — emmanuelvivier · 2026-07-26
- Silicon Valley is split on Chinese open-weight models now rivaling top U.S. systems — emmanuelvivier · 2026-07-26
- Kimi-K3 claims a perfect 6/6 on IMO 2026 Lean 4 proofs — songhan_mit · 2026-07-26