Researchers Accuse SOOFI Model of Benchmark Cheating

JJitsev · x · 2026-07-15

AI researcher Janus Leufer (@JJitsev) pointed out that the newly released SOOFI model allegedly manipulated benchmark data when claiming to match or outperform NVIDIA's Nemotron-3-Nano.

Comparative data shows that the Nemotron-3-Nano scores cited in SOOFI's report were severely understated (e.g., reporting 51.6 for MMLU-Pro versus NVIDIA's official 65.05). This "leaderboard brushing" tactic—deliberately lowering reference model scores to artificially boost their own model's performance—has sparked discussions about evaluation fairness in the AI community.

Related event: SOOFI benchmark claims challenged over leakage and baseline reporting(13 posts)→

Original post →

More from Models

Models channel →