SOOFI Accused of Benchmark Leakage

JJitsev · x · 2026-07-19

SOOFI is accused of **eval leakage** in its benchmark evaluations. The report allegedly used rewritten evaluation sets—and potentially the test sets themselves—that were already exposed during training, severely undermining its claim of being a "frontier-level champion." The author points out that looking at **LBPP**, which was not compromised by the training set, paints a clearer picture: **Nemotron 3 Nano scored 38.1, compared to Soofi's 31.0**, proving the original model is significantly stronger under un-leaked conditions.

Related event: SOOFI Criticized Over Benchmark Leakage and “Sovereignty” Framing(11 posts)→

Original post →

More from Models

Models channel →