DeepSeek V4 hits 67 t/s on dual GX10 GPUs
koalfied-coder · reddit · 2026-08-30
User reports achieving over 65 tokens/s sustained inference with DeepSeek V4 on dual GX10 GPUs, highlighting the utility of the 2570 context length evaluation.
More from Infra
- Bot Mesh: A social network with identity and payments for AI agents — Daniel_Farinax · 2026-08-30
- User Switches to Local Qwen 3.8 27B for Coding to Save API Costs — 4310sy · 2026-08-30
- Bezalel Offers Integrated Super Powers for AI Agents — Rasmic · 2026-08-30
- 19 General Latency Optimization Patterns for Faster AI Applications — blaizedsouza · 2026-08-30
- Superwall's side project policy leads to creation of open-source observability platform Maple — JordanMorgan10 · 2026-08-30
- Heterogeneous GPU benchmark of Qwen3.8-27B: eGPU layer-split and MTP acceleration analyzed — CoffeeToCode99 · 2026-08-30