Long-running benchmarks find Strata inference server failing full-build scenarios
julianharris · x · 2026-10-06
julianharris runs end-to-end benchmarks where models implement entire features from detailed specs — dozens of hours per run, 5 runs each for statistics, with 50+ completed "Miro clone MVP" builds logged across M5 Max, AMD Strix, Swift, MTPLX and gufo.
Strata, a promising new inference server for jamming large models into limited memory, passes all his smoke tests but bombs on full-build scenarios; he's axing it after n=1. The same benchmark recently surfaced important bugs in gufo and MTPLX. Root-cause analysis to follow on his blog.
Related event: Strata Inference Server Passes Smoke Tests but Fails Long-Run Benchmarks(2 posts)→
More from coding & agent
- Replit CEO: someone left an AI agent running overnight and woke up to $10,000 of wasted tokens — amasad · 2026-10-06
- Replit's Shlomi Fruchter: MCP, Harness and Skills are absurd concepts doomed like prompt engineering — shlomifruchter · 2026-10-06
- One Prompt Dump, Full App: Developer Wowed by Spawn Agent's Single-Shot Build — TAbrodi · 2026-10-06
- After Five Years, MATHPETS Launches as a Language and IDE for Agent-Based Models — jessi_cata · 2026-10-06
- Reading the source of 7 LLM eval tools uncovered 13 scoring bugs, 6 fixes merged — maverick_man1111 · 2026-10-06
- 6 open-source agent harnesses built for teams, from Deep Agents to Omnigent — femke_plantinga · 2026-10-06