Long-running benchmarks find Strata inference server failing full-build scenarios

julianharris · x · 2026-10-06

julianharris runs end-to-end benchmarks where models implement entire features from detailed specs — dozens of hours per run, 5 runs each for statistics, with 50+ completed "Miro clone MVP" builds logged across M5 Max, AMD Strix, Swift, MTPLX and gufo.

Strata, a promising new inference server for jamming large models into limited memory, passes all his smoke tests but bombs on full-build scenarios; he's axing it after n=1. The same benchmark recently surfaced important bugs in gufo and MTPLX. Root-cause analysis to follow on his blog.

Related event: Strata Inference Server Passes Smoke Tests but Fails Long-Run Benchmarks(2 posts)→

Original post →

More from coding & agent

coding & agent channel →