DIY claude-local: orchestrating local LLMs inside Claude Code with llama-swap
Distinct-Pie2389 · reddit · 2026-09-26
A detailed breakdown of 'claude-local', a custom hybrid orchestration system: an Opus-driven Claude Code setup that routes tasks to local models via an MCP server and skills, delegates cloud work to GLM/DeepSeek/Codex, and spawns pruned local subagents through llama-swap. Includes a precise explanation of how two parallel subagent sessions interleave on a single GPU slot with --parallel 1.
More from coding & agent
- Demo shows humans taking over AI agents mid-task on a real computer — aniketmaurya · 2026-09-26
- You.com uses OpenRouter as reranking layer: 3x fewer tokens, 84% benchmark accuracy — RichardSocher · 2026-09-26
- Microsoft standardizes on GitHub Copilot SDK across products, ex-engineer reveals — martinwoodward · 2026-09-26
- Creator builds 3D model in ComfyUI, audio-reactive stage with Sentinel — PurzBeats · 2026-09-26
- OOD Labs launches Sentinel, a $199.99 real-time node graph built for AI agents — PurzBeats · 2026-09-26
- Dev builds Litmus, a personalized AI-writing detector, tests 7 models — Sol sounds most human — chaseleantj · 2026-09-26