Strata on a single RTX 3090: 256k context at 38-61 t/s, 2x faster than llama.cpp

cezarducatti · reddit · 2026-10-03

A non-programmer Redditor shares a full writeup of compiling and running the Strata inference engine on a single RTX 3090 with 128GB RAM, far outperforming llama.cpp master: 1,650 t/s prompt processing (vs 700), 38 t/s generation at 182k context and 61 t/s short-context (vs 23), 256k fp16 KV context, 76% expert cache hit and MTP accept rates, and error-free tool calling in OpenCode. Key tweaks: Unsloth UD-Q3KXL quant with --compat-bf16, rebuilding for sm86 with MMQ enabled (2x Q4 prompt speed), disabling crash-prone fused kernels, moving the vision encoder to GPU, and per-model calibration with persisted expert cache. He found Strata's default recommended quant faster but lower quality, preferring Unsloth's for accuracy.

Original post →

More from Infra

Infra channel →