Nemotron 3 Nano Beats Soofi on LBPP
JJitsev · x · 2026-07-16
This discussion focuses on the LBPP benchmark results. The key point is that this benchmark falls outside the training contamination scope of Soofi's "rewritten benchmark set," making the results more reliable.
The conclusion drawn is that Nemotron 3 Nano significantly outperforms Soofi on this uncontaminated benchmark, scoring 38.1 vs 31.0. Replies also point out that signs of this performance drop without using rewritten benchmark training data were already visible in Soofi's own report.
Related event: SOOFI benchmark claims challenged over leakage and baseline reporting(13 posts)→
More from Models
- Unreleased 'GPT 6 Sol' model spotted in OpenAI's API — SteveEricJordan · 2026-09-11
- DeepSeek update keeps cache hits mid-conversation, cuts costs 36.6% — teortaxesTex · 2026-09-11
- AI Sextet offers 6 models free and unlimited for 14 days, including DeepSeek and Qwen — airesearch12 · 2026-09-11
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11