A local/remote LLM router dies as new model releases outpace its training
gaviniboom · reddit · 2026-10-03
Reddit user gaviniboom trained a 4B router model to split traffic between DeepSeek v4 Flash 0731 and GLM 5.2, matching GLM 5.2 performance locally at roughly equal token costs via OpenRouter, with plans to offload the DeepSeek side to a local server.
But when GLM 5.3 shipped, the router failed to keep up — performing worse than GLM 5.3 Flash — killing the project. The author says the code is research-grade and open to sharing or tuning advice. Key lesson: routers trained against specific model pairs go stale fast as upstream models iterate.
More from coding & agent
- User tasks Muse agent with filing a small claims court lawsuit — Scobleizer · 2026-10-03
- Grok Bot, Meta Muse, Dot, Hermes: which AI agent to use for which task — HarperSCarroll · 2026-10-03
- Stanford CS224V Agentic AI course opens: a bottom-up anti-hallucination stack for trustworthy agents — stanfordnlp · 2026-10-03
- Solo dev builds FTL-like airship game in a week with Claude, shipping 7 languages and 100 achievements — MosskeepForest · 2026-10-03
- Cloudflare's Clef decision models (27B and 9B) now available on Ollama — ollama · 2026-10-03
- OpenDots hits 1.5k GitHub stars as a self-hostable AI coworker template — aigclink · 2026-10-03