Grok 4.5 matches DeepSeek-V4-Pro on WeirdML but still misses the task format

teortaxesTex · x · 2026-07-23

Grok 4.5 is being compared against DeepSeek-V4-Pro on WeirdML, and the takeaway is that it looks roughly competitive on score but still struggles with the evaluation setup itself.

The quoted breakdown says Grok 4.5 (high) scored 46.4%, down from 49.9% for Grok 4.3. The model seems smarter overall, but it often fails to adapt to the task format, sometimes exploring the data instead of actually producing predictions, and in two tasks it ended up submitting nothing in any of the five iterations.

The broader point: WeirdML appears to reward models that have been specifically trained to respect the constraints of the benchmark, not just models that are generally capable.

Original post →

More from Models

Models channel →