Nemotron 3 Nano Beats Soofi on LBPP
JJitsev · x · 2026-07-16
This discussion focuses on the LBPP benchmark results. The key point is that this benchmark falls outside the training contamination scope of Soofi's "rewritten benchmark set," making the results more reliable.
The conclusion drawn is that Nemotron 3 Nano significantly outperforms Soofi on this uncontaminated benchmark, scoring 38.1 vs 31.0. Replies also point out that signs of this performance drop without using rewritten benchmark training data were already visible in Soofi's own report.
Related event: SOOFI benchmark claims challenged over leakage and baseline reporting(13 posts)→
More from Models
- Moonshot pauses Kimi K3 signups five days after launch as GPU demand surges — eyishazyer · 2026-07-21
- AI Diplomacy demo makes agents negotiate, ally, and betray each other — jamdac · 2026-07-21
- Newer models need a different prompting style, and old tricks can make outputs worse — emollick · 2026-07-21
- GLM-5.5 is said to arrive in 4 weeks with open weights — tanay_mehta · 2026-07-21
- Fable 5 is credited with a 3-variable counterexample to the Jacobian conjecture — Various-Affect4841 · 2026-07-21
- Ben’s Bites roundup highlights Kimi K3, Fable 5, Cursor costs and self-driving companies — Ben's Bites · 2026-07-21