Calling Out Score Manipulation in Benchmarking

JJitsev · x · 2026-07-15

The post mocks a deceptive "benchmarking 2" tactic: artificially lowering the evaluation scores of reference models to make one's own model look like a "frontier champion."

As an example, the author notes that while SOOFI S simply retrained the already open-source Nemotron-3-Nano, the report lists Nemotron-3-Nano scores significantly lower than those in the original report, thereby skewing the comparison.

Related event: SOOFI benchmark claims challenged over leakage and baseline reporting(13 posts)→

Original post →

More from Models

Models channel →