Rant: Gaming AI Model Benchmarks with Custom Metrics

JJitsev · x · 2026-07-13

The author rants about the current chaos in AI model evaluation: making one's own model look the strongest by inventing scoring metrics that no one else uses. For example, in the SOOFI evaluation, using a so-called "capability index" to evaluate base models made Qwen 3 32B appear weaker than practically insignificant models like Teuken.

Related event: Controversy Over Nemotron Reproduction and Custom Benchmarks(4 posts)→

Original post →

More from Models

Models channel →