DeepSeek-V4 Pro and Fable show lowest task correlation
zainhas · x · 2026-08-15
A comparison of DeepSeek-V4 Pro, Fable, and Sol reveals their behavioral similarities on tasks. DeepSeek-V4 Pro and Fable are the most diverse pair with a 0.39 correlation in outcomes, while Pro and Sol are the most alike at 0.54. The combination of all three models failed to solve only 4 tasks.
More from Models
- Anthropic's next model won't ship externally; closed-source AI progress said to be paused — bindureddy · 2026-08-15
- Zhipu AI Releases GLM-5.3: 743B Base, Focus on Coding and Cyber Defense — max_paperclips · 2026-08-15
- DeepSeek-V4 Pro benchmark: Top performance at ultra-low cost — zainhas · 2026-08-15
- DeepSeek-V4 Pro coding analysis: Stable but weaker on Rust — zainhas · 2026-08-15
- DeepSeek-V4 Pro beats Sol and Fable in coding tasks — zainhas · 2026-08-15
- Qwen3.8-Max launches on Together AI with 2.4T parameters — togethercompute · 2026-08-15