Grok 4.5 Likely Stronger Despite Sandbagging

teortaxesTex · x · 2026-07-12

The author references a validation data test used in **population genetics**, suggesting Sasha can be trusted to build a **rigorous validation dataset**. Their personal take: **Grok 4.5 is indeed very strong**, but the underperformance of Sol and Opus is likely due to **sandbagging (deliberately lowering performance)**, joking that Elon doesn't mind this behavior.

Original post →

More from Models

Models channel →