Kimi K3 (2.8T) runs at 1 token/s on a MacBook Pro, streamed from four SSDs
Argonautlabs · hn · 2026-09-09
An HN post showcases Deltafin, an open-source project (argonautlabsai/deltafin) that streams the 2.8T-parameter Kimi K3 weights from four SSDs and runs it on a MacBook Pro at roughly 1 token/s.
- The idea: weights don't need to fit in memory; they're streamed on demand from SSD.
- Speed is impractical, but it's an interesting feasibility demo for pushing local inference beyond RAM limits.
More from Infra
- Epoch AI: GPT's Quadratic Latency vs Claude's Linear May Explain Pricing Gap — scaling01 · 2026-09-09
- What 100 GW of compute really means: 876 TWh a year and a country-scale power system — shyamalanadkat · 2026-09-09
- vLLM's hard-won lessons: pipeline parallelism falters on warm agent turns, 2.7x decode on Kimi K3 — vllm_project · 2026-09-09
- vLLM details full-stack optimizations for real-world agentic serving on AgentX benchmark — vllm_project · 2026-09-09
- Google Cloud CEO: TPU servers pay back in ~1 year, half that of GPU servers — matt_slotnick · 2026-09-09
- DeepSeek v4.1 Flash flash sale: 58M tokens for $1, available for 2 days only — MicahBerkley · 2026-09-09