DeepSeek v4 Flash Sets New Cost/Accuracy SOTA on WeirdML
teortaxesTex · x · 2026-08-03
DeepSeek v4 Flash 0731 versions achieved 57.1% and 63.0% on WeirdML, setting a new cost/accuracy SOTA ahead of the recently repriced GPT 5.6 Luna.
Although the scores seem slightly below predictions, reviewers note the model often uses submissions just for data exploration rather than pure scoring. This suggests its true agentic capabilities are likely much higher, potentially reaching GLM 5.2 levels in the right harness.
More from Models
- Benchmarking 4 AI Models on Video Coding: Opus 5 Wins but Takes 5 Hours — brackify_gg_ · 2026-08-03
- Alibaba Launches 2.4T-Parameter Qwen3.8-Max, Open Weights Next Week — Deep_Ladder_4679 · 2026-08-03
- Claude Code Dominates Hugging Face Traffic, Driving Majority of AI Coding Agent Activity — vanstriendaniel · 2026-08-03
- Qwen 3.8 Max Outperforms Fable 5 in 3D Physics Scene Generation at 1/7 the Cost — rohanpaul_ai · 2026-08-03
- OpenAI's New Astra Model, AI Agents Escaping Sandboxes, and Pacing Calls — EverydayAI_ · 2026-08-03
- llama.cpp Adds MTP Support for Qwen3-Next, Enabling Full-Speed Inference — jacek2023 · 2026-08-03