Local DeepSeek V4 Flash Benchmark: 24 tok/s on Epyc + RTX 5090
IntravenusDeMilo · reddit · 2026-08-24
An author shared benchmarks running DeepSeek V4 Flash (UD-Q8KXL) on a self-built "low-rent" local inference machine (AMD Epyc 7663 + 256GB ECC DDR4 + RTX 5090 32GB).
- Performance: Achieved 23.8-24.6 tokens/sec on tasks with 100-128k context.
- Latency: Prompt processing (PP) ranged from 60ms to 385ms (likely due to caching).
- Comparison: Performance improved after removing DFlash and was faster than a temporary RTX 3090 setup.
The author expressed surprise that a model of this quality runs well locally without a $10k investment.
More from Embodied
- Tesla FSD talk reveals end-to-end training works like the human brain learning to drive — PTrubey · 2026-08-24
- China's robot demos tout dexterity, America's tout speed — the promo gap persists — tszzl · 2026-08-24
- Big Tech pushes AI wearables, sparking privacy and stalkerware fears in Europe — nordicinst · 2026-08-24
- Robots Learn to Taunt Opponents in RL — geoffwolfe · 2026-08-24
- Unitree Humanoid Outruns Bolt, OpenAI Pauses Frontier Training — PeterDiamandis · 2026-08-24
- Tsinghua Team Publishes in Science Robotics: Humanoid Robots Master Reactive Soccer Skills — 机器之心 · 2026-08-24