Researchers Accuse SOOFI Model of Benchmark Cheating
JJitsev · x · 2026-07-15
AI researcher Janus Leufer (@JJitsev) pointed out that the newly released SOOFI model allegedly manipulated benchmark data when claiming to match or outperform NVIDIA's Nemotron-3-Nano.
Comparative data shows that the Nemotron-3-Nano scores cited in SOOFI's report were severely understated (e.g., reporting 51.6 for MMLU-Pro versus NVIDIA's official 65.05). This "leaderboard brushing" tactic—deliberately lowering reference model scores to artificially boost their own model's performance—has sparked discussions about evaluation fairness in the AI community.
Related event: SOOFI benchmark claims challenged over leakage and baseline reporting(13 posts)→
More from Models
- AI Sextet offers 6 models free and unlimited for 14 days, including DeepSeek and Qwen — airesearch12 · 2026-09-11
- Anthropic publishes its most detailed threat report, including an AI-designed drone swarm case — soumitrashukla9 · 2026-09-11
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11