Strata on a single RTX 3090: 256k context at 38-61 t/s, 2x faster than llama.cpp
cezarducatti · reddit · 2026-10-03
A non-programmer Redditor shares a full writeup of compiling and running the Strata inference engine on a single RTX 3090 with 128GB RAM, far outperforming llama.cpp master: 1,650 t/s prompt processing (vs 700), 38 t/s generation at 182k context and 61 t/s short-context (vs 23), 256k fp16 KV context, 76% expert cache hit and MTP accept rates, and error-free tool calling in OpenCode. Key tweaks: Unsloth UD-Q3KXL quant with --compat-bf16, rebuilding for sm86 with MMQ enabled (2x Q4 prompt speed), disabling crash-prone fused kernels, moving the vision encoder to GPU, and per-model calibration with persisted expert cache. He found Strata's default recommended quant faster but lower quality, preferring Unsloth's for accuracy.
More from Infra
- DwarfStar 4 (ds4) lets you run DeepSeek V4.1, Qwen and GLM locally — yogthos · 2026-10-03
- Local LLM users question GPU upgrades as prices outpace performance gains — masiha97 · 2026-10-03
- Fireworks launches cache-aware FireRouter with Opus: coding costs cut 57% at 98% accuracy — Madisonkanna · 2026-10-03
- OpenAI reportedly weighed $100M Hugging Face investment before Nvidia deal, with chip distribution in play — VraserX · 2026-10-03
- MegaCapybara: RTX 5090-only inference engine hits 2000+ t/s, 2x faster than Ninfer — BringTea_666 · 2026-10-03
- "Opposing cheaper electricity because data centers benefit" sparks AI power-policy fight — 2C_ornot2C · 2026-10-03