Pure-CPU DeepSeek V4.1 Inference Project Hits 6 TPS on a Xeon for Overnight Agent Jobs
Qwen30bEnjoyer · reddit · 2026-09-11
A Reddit user open-sourced Day1DeepseekV4.1-CPU, a hacky setup running DeepSeek V4.1 entirely on a lab Xeon server — 30 TPS prefill and 6 TPS decode with n-gram table offload on 50% of threads. The idea: unlimited slow tokens for overnight or multi-day agentic jobs, with a watcher script that kills the process within 15 seconds if someone else needs the machine for genome assemblies. The code was vibecoded by Opus 5.0 and unreviewed; the author hopes kernel experts will contribute PRs or forks.
More from Infra
- 1:26 continuous aerial AI video made entirely on a Mac with MiniMax H3 — cocktailpeanut · 2026-09-11
- KV cache gets QAT too: why this model beats others at fp4 KV cache — stochasticchasm · 2026-09-11
- Commentary: Anthropic loads shift to Google plus AWS slice, OpenAI doubles down on Azure — ericwdolan · 2026-09-11
- Nvidia claims Vera Rubin delivers 50X throughput per MW and 35X lower token cost vs Blackwell Ultra — Beth_Kindig · 2026-09-11
- Epoch AI: GPT long-context latency scales quadratically, matching price jumps — Jsevillamol · 2026-09-11
- RTX 3090 mini-bench: ninfer cuts TTFT from 3.4s to 29ms, prompt processing ~76x faster — milkipedia · 2026-09-11