Grok 4.5 Likely Stronger Despite Sandbagging

teortaxesTex · x · 2026-07-12

The author references a validation data test used in population genetics, suggesting Sasha can be trusted to build a rigorous validation dataset.

Their personal take: Grok 4.5 is indeed very strong, but the underperformance of Sol and Opus is likely due to sandbagging (deliberately lowering performance), joking that Elon doesn't mind this behavior.

Original post →

More from Models

Models channel →