Lucebox Zero 495 pairs Ryzen AI MAX+ PRO 495 with R9700 for 50 tok/s local LLMs
samsja19 · x · 2026-10-03
Lucebox Zero 495 is now open for pre-order, pairing AMD's Ryzen AI MAX+ PRO 495 (up to 192GB unified memory) with a 32GB Radeon AI PRO R9700 in one box.
The key is its inference engine's split strategy: dense layers, the hottest experts, and the drafter go on the R9700 since nearly every token reads them, while the long tail of experts stays in the 495's unified memory — both chips compute simultaneously. Throughput jumps from 15 tok/s on the 495 alone to 50 tok/s with the R9700.
Options include 4TB SSD and a 100 GbE cluster kit; the first 200 units get $1,000 off, with January delivery expected.
More from Infra
- Running 256k-context open models on 2x RTX 3090 for months: a home server LLM retrospective — knighty1981 · 2026-10-03
- Prime Intellect compresses MLA KV cache in NVFP4, fitting ~50% more tokens than FP8 — TheZachMueller · 2026-10-03
- Runware launches Serverless GPUs: $0 while idle, from $0.63/GPU-hour — aziz4ai · 2026-10-03
- Burkov: Generative AI Only Makes Money for GPU Sellers, Echoing Dotcom Bubble — burkov · 2026-10-03
- Deriving KV-cache placement from abstract representations: prefill and inference are linked — vtabbott_ · 2026-10-03
- NVIDIA long stopped just selling GPUs: from CUDA to sovereign, agentic and physical AI — sudoraohacker · 2026-10-03