SOOFI benchmark claims challenged over leakage and baseline reporting

On July 15–16, criticism centered on how SOOFI evaluated and marketed its model. Janus Leufer (@JJitsev) argued that claims such as matching or surpassing NVIDIA’s Nemotron-3-Nano, and even being the “strongest open model,” are not well supported because the comparison may be contaminated by rewritten benchmark data and by weak baseline reporting. The issue matters because it touches a core trust problem in open releases: overlap between training data and test sets can make benchmark gains look like capability gains.

Core criticisms

According to @JJitsev, SOOFI’s training data included rewritten versions of evaluation sets, including GPQA Diamond and some German benchmarks. He contrasted this with version details around Nemotron: in his account, the original Nemotron 3 Nano was not trained on those rewritten benchmarks, while Nemotron 3 Super was. He therefore argued that using Nemotron-3-Nano as the headline comparison mixes real model ability with benchmark contamination. Separately, @teortaxesTex used GPQA as an example to argue that even lightly rewritten test items can still heavily pollute evaluation if they enter training, noting a case described as training for 10 epochs.

Which comparisons are more informative

@JJitsev highlighted LBPP in SOOFI’s Table 5 as more useful because, in his telling, it was not part of the rewritten benchmark set. On that benchmark, he said Nemotron 3 Nano clearly outperformed SOOFI. He also disputed SOOFI’s stronger-German narrative: SOOFI already had more German training data, and if rewritten German benchmarks were also included in training, then the resulting comparison would be even less reliable.

Score reporting and wider fallout

Another part of @JJitsev’s critique was that SOOFI’s report presented Nemotron-3-Nano scores lower than NVIDIA’s original public results, which he said made SOOFI S look more “frontier-level” than a fair comparison would justify. Based on the combination of low baseline reporting and benchmark leakage concerns, he said there is currently no evidence that SOOFI matches its source model Nemotron 3 and urged the project to stop overclaiming. @eliebakouch extended the criticism by questioning SOOFI’s “sovereign” positioning, arguing that the model appeared very close to Nemotron-3-Nano in architecture and shared about 80% of the data mixture, while pointing to the same GPQA Diamond-style contamination concern.

2026-07-15 ~ 2026-07-16 · 13 related posts

1 near-duplicate retellings: JJitsev