antirez runs DeepSeek v4.1 Flash locally on a 128GB M5 Max, SSD streaming surprisingly fast
antirez · x · 2026-09-11
antirez demonstrated DeepSeek v4.1 Flash running locally via DwarfStar on a 128GB M5 Max, with SSD-based expert streaming delivering surprisingly fast inference. He credits recent SSD streaming changes that retain the right experts, and possibly DS4.1 reusing the same experts more. He plans to open QA online when ready.
More from Infra
- OpenAI Agents API hits public beta; Cloudflare ships sandbox integration for cloud Codex agents — ritakozlov · 2026-09-11
- DeepSeek-V4.1-Flash lands on Fireworks: 552B MoE for coding and agents at 1/40th claimed cost — lqiao · 2026-09-11
- OpenAI Agents API Meets Cloudflare: Deploy a Logged Agent in 4 Minutes — craigsdennis · 2026-09-11
- Training a 6-Expert MoE GPT-2 From Scratch on a Single RTX 3090 in 8 Days — rasbt · 2026-09-11
- B3IQ Sells Eight Figures of GPUs in Two Weeks, Bets AI Infra Is a $100B Market — templecrash · 2026-09-11
- It Cost $100 in API Credits for an AI Agent to Install Free Software — MartinGTobias · 2026-09-11