Strata calibrate nearly tripled decode speed: 256K context on a 16GB GPU
MoonsvnLyn · reddit · 2026-10-04
On a 5070 Ti (16GB) + 14700KF + 96GB RAM, the author ran Qwen3.8-Flash-Next IQ3S with 262K context through Strata (KV cache in RAM streaming a 32K window into VRAM), lifting decode from 17.2 tok/s to 43 cold and 53.5 with prefix cache — with full repro steps:
- --pool-workers 13 instead of 19 (+47% alone): missed experts are computed on CPU and every verify window waits for the slowest thread, so E-cores dragged everything down; defaults were tuned on a 6-core Ryzen with no E-cores. Intel 12th–14th gen hybrid chips should run calibrate.
- --spec 6 with --spec-min-p 0.70: deeper speculative windows only extended when confident; acceptance rose from 53% to 70–79%, making deep windows free speed.
- --pcie-frag 0: on PCIe 5 x16 with a fast CPU, computing a missed expert in place beats copying it over the bus (calibrator curve: 42.1 tok/s at 0.0 down to 20.7 at 0.75).
Caveat: re-running setup rewrites the config and drops profile flags. Full numbers live in the author's tuning doc on the Strata repo, plus an upstream issue filed for the unmeasured IQ3S@262K data point.
More from coding & agent
- Local LLM benchmarks are mostly noise: c=1 tokens/s hides real concurrency performance — TheZachMueller · 2026-10-04
- Even Opus 5.5 ships vulnerabilities when your vibe coding requirements are vague — gefei55 · 2026-10-04
- Too many agents to track: David Khourshid loses track of what his own bots do — DavidKPiano · 2026-10-04
- Harness engineering reduced to prompts and tools, but it's what makes agents reliable — techNmak · 2026-10-04
- A model alone isn't an agent: a first-principles handbook on agent harness engineering — techNmak · 2026-10-04
- Building a Miro clone 5x on 3 local rigs: tokens/sec is useless, thinking variance hits 5x — julianharris · 2026-10-04