Strata on a 4090 48GB: conversation parking buys 35x, n-gram table in RAM just 1%
Shot-Ad-4147 · reddit · 2026-10-04
A detailed single-machine benchmark of Strata 0.1.38 on i9-13900K + RTX 4090 48GB + 128GB DDR5, using a real 91,836-token prompt.
Four-arm test of n-gram table placement
- Keeping the full 28.8 GB table in RAM (+28.4 GiB) buys only +0.65% warm prefill / +1.2% decode — the SSD-resident table isn't a bottleneck here.
- --ple-io mmap is worse: first long prompt takes 72.8s vs 19.5s (3.7x slower) and silently disables the row cache.
- The default 90 MB row cache already absorbs 92% of reusable traffic (3,114 MB → 242 MB on second pass).
Methodology lesson: the author's earlier +4.7% claim came from a wrong baseline — session-to-session spread on identical config was 11.5% (149.1 → 166.3 tok/s). A/B runs must happen back-to-back in the same session.
Biggest win — conversation parking (35x): with --conversation-cache-mib 8192 and 4 slots, returning to a 91,836-token conversation drops from 19.5s to 537 ms for just 3.1 GB of RAM.
Other wins: --prefill auto:32768 gives +9.7% prefill / +13.6% decode at 78.7K context; --calibrate found another +3.2% by changing CPU workers 23 → 12. Baseline throughput: 128-151 tok/s decode, 4,249-5,154 tok/s prefill.
More from Infra
- Cloudflare's open-weights decision model Clef scores all answers in one pass, 53ms per decision on RTX 5090 — michellechen · 2026-10-04
- Tencent to lease ~100k chips from Oracle for ~$7bn; BoE flags agent risk — MirelaXhota · 2026-10-04
- Qwen 27B on a single R9700: 3-bit rotation quant buys 569K-token cache and 12x faster agent turns — evp-cloud · 2026-10-04
- Strata runs Qwen 3.8 Flash Next (125B) on a single RTX 4090 at 100 tok/s — snehesht · 2026-10-04
- Cloudflare Workers tip: loop cron work back near your DB, p90 queries drop 75% — DanielLockyer · 2026-10-04
- Aletheia: AMD's $40B server CPU 2027 target supply-constrained, 70% growth seen in 2028 — zephyr_z9 · 2026-10-04