Researchers Accuse SOOFI Model of Benchmark Cheating
JJitsev · x · 2026-07-15
AI researcher Janus Leufer (@JJitsev) pointed out that the newly released SOOFI model allegedly manipulated benchmark data when claiming to match or outperform NVIDIA's Nemotron-3-Nano.
Comparative data shows that the Nemotron-3-Nano scores cited in SOOFI's report were severely understated (e.g., reporting 51.6 for MMLU-Pro versus NVIDIA's official 65.05). This "leaderboard brushing" tactic—deliberately lowering reference model scores to artificially boost their own model's performance—has sparked discussions about evaluation fairness in the AI community.
Related event: SOOFI benchmark claims challenged over leakage and baseline reporting(13 posts)→
More from Models
- Moonshot pauses Kimi K3 signups five days after launch as GPU demand surges — eyishazyer · 2026-07-21
- AI Diplomacy demo makes agents negotiate, ally, and betray each other — jamdac · 2026-07-21
- Newer models need a different prompting style, and old tricks can make outputs worse — emollick · 2026-07-21
- GLM-5.5 is said to arrive in 4 weeks with open weights — tanay_mehta · 2026-07-21
- Fable 5 is credited with a 3-variable counterexample to the Jacobian conjecture — Various-Affect4841 · 2026-07-21
- Ben’s Bites roundup highlights Kimi K3, Fable 5, Cursor costs and self-driving companies — Ben's Bites · 2026-07-21