Running an Opus-level coding agent locally at nearly 2x Claude Opus speed, for free
julianharris · x · 2026-09-11
Julian Harris's hands-on long read claims Aug 2026 marks the tipping point for local AI coding: he runs an Opus-level coding agent locally for free at nearly 2x Claude Opus speed.
Key details:
- Tensor parallelism via RDMA across two MacBooks (50% each) reaches 25 t/s, vs 15 t/s with SSD streaming on a single machine—likely optimizable much further
- Benchmarks show doubling context halves qwen-3.8-27b throughput, while qwen-3.8-flash-next is barely affected
- Motivation: Anthropic's erratic access limits, rate caps (5-hour windows), and censorship made the commercial relationship untenable
- He built Ceetrix, an MCP spec-governance server that keeps coding agents from "behaving like lazy, gaslighting interns"
Fully hand-written article with days of first-hand experimentation.
More from coding & agent
- Sierra Catalina proposes Context Layer: a six-step user-controlled protocol for sharing minimal context across AI agents — sierracatalina · 2026-09-11
- Replit-Databricks integration hits GA with native Lakebase support and AI schema approvals — amasad · 2026-09-11
- Descript criticized for offering no agentic API: editing locked behind its token-charging Underlord agent — heyneighbor · 2026-09-11
- Open-source voice agent demo works via web or phone, easy to customize in routes/agent.ts — juberti · 2026-09-11
- Mobile Easy Use: Open-Source MCP Tool Lets Coding Agents Debug Real Android/iOS Apps — James_CSH · 2026-09-11
- Cursor Origin repos can now deploy on Netlify, in beta — thisiskp_ · 2026-09-11