DeepSeek V4 Pro benchmarks close to Opus 5 on KernelBench-Hard
teortaxesTex · x · 2026-08-21
Benchmark results show DeepSeek V4 Pro achieving 9.35% of roofline performance on the TopK task in KernelBench-Hard, closely trailing Opus 5 at 9.46%. The data also compares performance across models like Kimi K3, Fable 5, and Qwen 3.8 Max on various tasks including FP8, KDA, and Paged attention.
More from Models
- Musk Confirms Work to Improve Grok's Writing Skills — mark_k · 2026-08-21
- Why 'Full Pass Rate' is a flawed metric for LLM evaluation — xeophon · 2026-08-21
- ARC Prize Adds Model Comparison, Gemini 3.7 Flash Scores High — mhmazur · 2026-08-21
- Anthropic's Fable Breaks RareBench Record After Relaxing Filters — danielmckinn0n · 2026-08-21
- NVIDIA Explains Omni-Models: Unified Architecture for Text, Images, Audio, Video, and Actions — NVIDIA Developer · 2026-08-21
- Monitors Detect Significant Behavior Shift in Claude Opus — altryne · 2026-08-21