Grok 4.5 matches DeepSeek-V4-Pro on WeirdML but still misses the task format
teortaxesTex · x · 2026-07-23
Grok 4.5 is being compared against DeepSeek-V4-Pro on WeirdML, and the takeaway is that it looks roughly competitive on score but still struggles with the evaluation setup itself.
The quoted breakdown says Grok 4.5 (high) scored 46.4%, down from 49.9% for Grok 4.3. The model seems smarter overall, but it often fails to adapt to the task format, sometimes exploring the data instead of actually producing predictions, and in two tasks it ended up submitting nothing in any of the five iterations.
The broader point: WeirdML appears to reward models that have been specifically trained to respect the constraints of the benchmark, not just models that are generally capable.
More from Models
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- Terminal Bench v4: GLM-5.3 Leads at 41.9%, Kimi-K3 Underwhelms at 12.6% — Ok_Warning2146 · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11