SOOFI Accused of Eval Leakage and Overhype

JJitsev · x · 2026-07-19

The author points out a clear eval leakage in SOOFI's comparisons: its cloned training set has already seen most of the evaluation data, and GPQA even includes the test set. Therefore, claiming it is a "frontier-level champion" is entirely baseless.

He further criticizes the narrative of the SOOFI report: it exaggerates its "transparent" and "sovereign" positioning while failing to clearly state that it was reproduced on Nemotron-3-Nano / the open Nemotron stack. It also uses a self-built "capability index" to support its comparative conclusions, showing obvious packaging issues.

Related event: European Open-Source Model SOOFI Accused of Eval Leakage and Overhype(11 posts)→

Original post →

More from Models

Models channel →