Building a Miro clone 5x on 3 local rigs: tokens/sec is useless, thinking variance hits 5x
julianharris · x · 2026-10-04
The author ran his "Miro MVP Clone" build on three local AI rigs (AMD 395, MacBook M5 Max, RTX 4090), five builds each, to measure stability.
Key findings:
- Timing variance exceeds 10%, while quality is more consistent (7%).
- "Tokens per second" is a useless metric: some models are extremely inefficient, others aren't, and a few show huge "thinking variance" — the same spec takes 10 minutes on one run and 50 minutes on an identical repeat (5x).
- The author has theories for why but doesn't elaborate in this post.
A rare hands-on measurement of AI coding agent run-to-run reliability, with practical implications for developers relying on agents in real engineering.
More from coding & agent
- xAI Employees Run 50+ Grok Bots, Orchestrated by Manager Bots — petergyang · 2026-10-05
- Paper Finds Personal Agents Get Worse as Memory Notes Pile Up — rohanpaul_ai · 2026-10-05
- Dev Discovers You Can Mirror a Mac-Running Simulator Onto Your Phone — itsOmSarraf_ · 2026-10-05
- Agent Memory Has a Sweet Spot: 10 Lines Optimal, Code Beats Rules for Tracking — rohanpaul_ai · 2026-10-05
- Raven V2 Brings Agentic Modeling to Rhino/Grasshopper, 15k+ Seats — burhop · 2026-10-05
- Indie Dev Ships Three AI-Made Games in Half a Month, First Already on App Store — ezshine · 2026-10-05