Hobbyist runs 512K context locally on CPU/RAM/SSD/GPU hybrid at 1,526 tok/s prefill
HankYeomans · x · 2026-10-08
- After much experimentation, the author broke through 512K token context locally on a "Frankenstein" hybrid rig combining CPU, RAM, SSD, and GPU, running a DeepSeek V4.1 Flash variant.
- Reported performance: prefill throughput of 1,526 tok/s at 16K/32K context—an interesting data point for feasibility of huge-context local inference.
More from Infra
- Genesis Mission partners pledge $2.4B in compute; NSF and DOE add $100M — AllThingsApx · 2026-10-08
- AWS CEO on rebuilding the cloud for agents: $220B 2026 CapEx, 2M NVIDIA GPUs ordered — a16z · 2026-10-08
- Strix Halo NPU finally put to work: local 125B MoE replaces 95% of cloud coding agent calls — stereohype · 2026-10-08
- Nvidia-backed data center firm's IPO demand plunges, exposing cracks in AI funding boom — SumitGup · 2026-10-08
- GlobalFoundries signs $2B deal to supply TSMC's CoWoS advanced packaging from US soil — pstAsiatech · 2026-10-08
- Microsoft goes all-in on local AI: hybrid intelligence Windows, 1.6-bit DeepSeek V4 Flash — Sam Witteveen · 2026-10-08