DeepSeek 284B on 4x3090: Prefill Hits 1906 tok/s After Optimization
max_paperclips · x · 2026-08-05
Developer @superalesha shared extreme optimization results running the DeepSeek-V4-Flash 284B model (2-bit quantized) on 4x RTX 3090 GPUs.
- Performance Leap: After removing bottlenecks between the CPU and GPUs, prefill speed skyrocketed from 433 tok/s to 1906 tok/s (a 4.3x increase at 98K context depth).
- Decode Untouched: Decode speed remained stable at 37-40 tok/s.
- Zero Weight Changes: The optimization was achieved in one day without altering a single model weight, purely by fixing hardware communication walls.
More from Infra
- Deleting 90% of Weights: Song Han's Journey to Efficient AI & Quantization — JafarNajafov · 2026-08-05
- Qwen3.6 35B NVFP4 Hits 3,000 Tokens/s on a Single RTX PRO 6000 — max_paperclips · 2026-08-05
- Cloudflare CEO Predicts Bot Traffic Will Reach 1,000x Human Traffic Within Five Years — 0xSammy · 2026-08-05
- Running Minimax H3 on RTX 3060 Takes 1.5 Hours for a 15s Clip, Dev Seeks Optimization — SMPTHEHEDGEHOG · 2026-08-05
- Can RTX 5060Ti 16GB Run Qwen 27B? Users Discuss Local LLM Hardware — Yanzihko · 2026-08-05
- AI Profits Drive US Stocks to Record Highs: Palantir Revenue Jumps 93% — nordicinst · 2026-08-05