Local DeepSeek prefill optimization doubles throughput to 1,820 tok/s in two days
HankYeomans · x · 2026-10-09
A developer shared progress on their self-assembled ("Frankenstein") local DeepSeek inference setup:
- Workload: 16K prefill. Two days ago: 860 tok/s, 18.6 s.
- After optimization: 1,820 tok/s, 8.8 s — throughput doubled and latency nearly halved within two days.
No specific method disclosed, but a notable speed reference point for anyone tuning local LLM deployments.
More from Infra
- $500/month API bills vs $14k local rig: Mac Studio 512GB or 2x DGX Spark? — rodrigodevbits · 2026-10-09
- Solari launches agent infrastructure: 8ms browsers, 10x faster than Browserbase — Scobleizer · 2026-10-09
- ARK analyst: the viral 'no high income, low energy country' chart is a snapshot — skorusARK · 2026-10-09
- Splash 1.3.0 cuts local agent first-token latency from 19s to 1s via SSD offloading on M6 Mac — BeidiChen · 2026-10-09
- Latency cut 2.7s to 0.4s on GLM via scheduling, no model or hardware changes — dbreunig · 2026-10-09
- fal and a16z host AI Infra Day on Oct 20 with BFL, Krea and ElevenLabs engineers — gorkem · 2026-10-09