Thoughtworks Engineer Spends 4 Weeks Testing Whether Local Models Are Viable for Coding
bibryam · x · 2026-10-04
Birgitta Böckeler, Distinguished Engineer at Thoughtworks, has published a memo assessing whether locally hosted open models are viable for coding — focusing on agentic coding rather than autocomplete.
Key points:
- After years of disappointment, she spent roughly 4 weeks re-testing local models, prompted by widespread claims that they've become genuinely good at coding.
- Hardware used: an Apple M3 Max (48GB RAM) and an Apple M5 Pro (64GB RAM).
- She evaluates not just raw capability but ease of use for developers who don't want to dig through specs and tooling.
- Many interacting factors make it hard to pick the best setup under resource constraints, and hard to separate signal from noise in people's online success stories.
- A counterintuitive finding: in her automated eval setup, one model clearly performed worse on the stronger machine.
This first memo covers the framework of viability factors; a follow-up will detail her hands-on experience.
More from coding & agent
- 8 Emerging Standards for AI Agents You Should Know — bibryam · 2026-10-04
- Developer Guesses Open Coding Models Are Less Than Six Months Behind Frontier — heyneighbor · 2026-10-04
- Harness vs. Loop Engineering: The Two Distinct Problems in Building Reliable AI Agents — goyalshaliniuk · 2026-10-04
- ffmpeg-skill turns coding agents into video editors: 42 local FFmpeg tools, auto-verification — udmrzn · 2026-10-04
- Strata calibrate nearly tripled decode speed: 256K context on a 16GB GPU — MoonsvnLyn · 2026-10-04
- One prompt freed 39.7 GB: Claude Code + ccmd MCP safely cleans dev caches — julsimon · 2026-10-04