A 397B MoE now runs on one RTX PRO 6000 with 96GB VRAM
mrstoatey · reddit · 2026-07-27
Krasis, a MoE-focused runtime for streaming very large models through limited VRAM, now runs Ornith-1.0-397B interactively on a single RTX PRO 6000 Blackwell 96GB.
What they measured
- 1,346 tok/s prefill at 10k tokens
- 2,354 tok/s prefill at 39,920 tokens
- 23.6 tok/s decode over 50 tokens
- 20.4 tok/s sustained decode over 250 tokens
- With Adaptive Cold Mass Pruning, decode improved to 25.7 tok/s while skipping only about 1.8% of routed probability mass on average
How it works
- Experts stay in CPU RAM and are dynamically brought into VRAM
- About 43% of routed experts were resident in VRAM in this run
- Peak process RAM was about 202GB, so the author says 256GB RAM is feasible on a consumer DDR5 motherboard
Additional notes
- Smaller MoEs that fit in system RAM can run much faster, e.g. 35B-class models at 117 tok/s decode on a 5090
- The same model can even run on a single RTX 5090 32GB at around 7.9 tok/s decode if you are patient
More from Infra
- LightRAG v1.5.3 adds a one-time Milvus migration and hardens production edge cases — JeremyCMorgan · 2026-07-28
- llama.cpp adds DSpark speculative decoding and asks for speed results — pmttyji · 2026-07-28
- Nvidia puts an open vision-language-action model on Hugging Face — theteknosaur · 2026-07-28
- “Model eats harness” is really about deployment-driven harness evolution — m_wulfmeier · 2026-07-28
- Agentic AI system reportedly cost $1.2M a month before being cut to $100K — emmanuelvivier · 2026-07-28
- Wistron’s Early Nvidia Bet Turned It Into One of AI’s Biggest Supply-Chain Winners — pstAsiatech · 2026-07-28