Dev pushes Qwen 27B to 70-90 tok/s on a 4090 in a local AI coding workflow
julianharris · x · 2026-09-20
Developer julianharris shares his fast local AI coding setup:
- Tuning win: a setting bumped 4090 + Qwen 27B from 50-70 to sustained 70-90 tok/s — roughly 2x Claude Opus throughput
- Stack: the minimal Pi agent harness plus an MCP extension wired into his spec management/governance system (Ceetrix)
- Case study: Qwen found a bug that looked like it required a rewrite, but Ceetrix task context revealed the feature existed — just behind a menu instead of a button
- Pi: supports skills, AGENTS.md, four modes (interactive/print/RPC/SDK), and self-customization via extensions
More from coding & agent
- TradingAgents: an open-source multi-agent LLM trading framework in Python — mdancho84 · 2026-09-20
- This guy used an AI agent to profile every eligible bachelor in the city for two cents — gregmushen · 2026-09-20
- OpenHarness: open-source workbench for orchestrating coding agents beyond code — dee_hw · 2026-09-20
- Lessons from a cited paper-writing LangGraph agent: token blowups, fake sources, and four fixes — Altruistic-Video-849 · 2026-09-20
- Plugin4Shell zero-click RCE hits Claude Code, Codex, Copilot and Gemini CLI days before NIST IR 8587, exposing the gap in agent authorization — docybo · 2026-09-20
- px0 editor ships git status streaming via SSE, checking just 3 files instead of polling — arpit_bhayani · 2026-09-20