ds4flash runs locally at 40 t/s single-thread, 170 t/s on six, all under 100W
l4rz · x · 2026-08-27
l4rz shares local inference results: a setup consuming under 100W delivers 40 tokens/s on a single request thread and about 170 t/s across six threads, calling it his "plan B".
More from Infra
- Nvidia projects 70% revenue growth for fiscal 2028, beating analyst expectations of 44% — firstadopter · 2026-08-27
- Nvidia's NVHBM Brings 30% More Bandwidth to NVLink Fusion; Amazon Annapurna First Partner — nordicinst · 2026-08-27
- NVIDIA CFO forecasts 70% growth next year, targeting ~$700B revenue — BenBajarin · 2026-08-27
- AWS and NVIDIA to Deploy 2 Million Additional GPUs for Agentic and Physical AI — nvidia · 2026-08-27
- Gemini 3.5 Transcribe Now Available on Vercel AI Gateway — osanseviero · 2026-08-27
- Nvidia guides $108B Q3 revenue, doubling growth even without China data center sales — inductionheads · 2026-08-27