Users Urge Neoclouds to Boost DeepSeek Speeds to 300-400 tok/s for Fast Models
brandon_galang · x · 2026-08-05
The author observes that despite previous competition driven by Kimi K3 to push token speeds, current inference speeds for new models on neoclouds remain underwhelming.
They urge providers to crank speeds up to 300-400 tok/s, creating an ideal "fast general model" for tasks that don't require maximum intelligence.
More from Infra
- Europe Pledges €30B for AI Gigafactories, Only €1B Actually Committed — sanjaykalra · 2026-08-05
- Modal Optimizes Serverless Architecture to Reduce Network Latency — AAAzzam · 2026-08-05
- Elon Musk: AI Memory Demand Growing Over 200% Annually, Supply Lagging — firstadopter · 2026-08-05
- Raspberry Pi + Hailo 10H Powers Real-Time Local Voice Assistant — martincerven · 2026-08-05
- Ex-Jane Street Exec Explains How They Cut Trading System Latency by 100x — ivan_bezdomny · 2026-08-05
- Musk: AI and Robotics Will Trigger Unprecedented Global Bandwidth Demand — XFreeze · 2026-08-05