Grok 4.5 matches DeepSeek-V4-Pro on WeirdML but still misses the task format
teortaxesTex · x · 2026-07-23
Grok 4.5 is being compared against DeepSeek-V4-Pro on WeirdML, and the takeaway is that it looks roughly competitive on score but still struggles with the evaluation setup itself.
The quoted breakdown says Grok 4.5 (high) scored 46.4%, down from 49.9% for Grok 4.3. The model seems smarter overall, but it often fails to adapt to the task format, sometimes exploring the data instead of actually producing predictions, and in two tasks it ended up submitting nothing in any of the five iterations.
The broader point: WeirdML appears to reward models that have been specifically trained to respect the constraints of the benchmark, not just models that are generally capable.
More from Models
- Repligate says Claude Opus 3 appears to evolve without changing its weights — repligate · 2026-07-27
- “Opus 5” post lands as a rebenchmarking-at-scale AI joke — kalomaze · 2026-07-27
- Top models now write worse than a year ago, critic says — dbreunig · 2026-07-27
- MPT-30B radar charts became an unexpectedly controversial design choice — code_star · 2026-07-27
- Local Gemma 4 31B starts acting sarcastic and users cannot reproduce it — n0head_r · 2026-07-27
- Google’s Gemini 3.6 Flash could win by matching Sonnet quality at a lower cost — haider1 · 2026-07-27