Passes smoke tests, fails full builds: long-running benchmarks doom Strata inference server
julianharris · x · 2026-10-06
Developer julianharris tested Strata, a promising new inference server for running large models on limited memory. It passed all his smoke tests but bombed on his end-to-end "full build" scenarios, so he is dropping it after n=1.
His benchmarking is unusually rigorous: models implement entire features from extremely detailed specs (tasks plus tests), he has completed a "Miro clone MVP" over 50 times, each run takes dozens of hours, and he repeats runs five times for statistical validity — with all conversations and logs stored in a database across M5 Max and AMD Strix hardware and software stacks like Swift, MTPLX, and gufo. He notes the same benchmarks surfaced important bugs in gufo and MTPLX last week; detailed Strata findings to follow.
Related event: Strata Inference Server Passes Smoke Tests but Fails Long-Run Benchmarks(2 posts)→
More from Infra
- PlanetScale Engineer Explains Kubernetes Feedback Loops by Running Postgres by Hand — bibryam · 2026-10-06
- One ENV Var Cut This Cloud Run Job's Billable Time by ~50%: NODE_COMPILE_CACHE — TechNadu · 2026-10-06
- 8GB RTX 4060 Ti Tuned for 2x-9x Faster Local LLM Inference vs llama.cpp Defaults — ExxploreCraft · 2026-10-06
- Dual R9700 RDNA4 Setup Hits 5000 tok/s Prefill on Qwen 3.8 27B, Seeks Better Engines — N34257 · 2026-10-06
- Why Don't We Build Computer Networks Like Living Things? — Sanity · 2026-10-06
- Google's AI infra chief: at 100K accelerators, FLOPS is a vanity metric — goodput is what matters — Training Data (Sequoia) · 2026-10-06