El Pulpo 0.1.0: A 12.7MB Proxy and Load Balancer for Local LLM Inference
zaytzev · reddit · 2026-09-16
A developer released El Pulpo 0.1.0, an open-source proxy/load balancer for local LLM inference, in a 12.7MB image. It unifies model provider settings across agents (no more reconfiguring when switching models or networks like Tailscale), adds token usage monitoring to see if devices pay for themselves, and load-balances across instances — built around the author's Qwen3.8 Flash Next setup.
More from Infra
- Which 'token brokers' give back to open source? New data ranks upstreamed PRs to OSS inference engines — michellechen · 2026-09-17
- XeBoostLM: native C++ local LLMs on Intel NPUs and iGPUs, zero Python — Spiritual-Ad-5916 · 2026-09-17
- Chips, batteries, motors fell 99%+ in 34 years — ARK says AI is now deflating 99%+ annually — skorusARK · 2026-09-17
- MLPerf Inference v6.1 draws record 30 submitters, adds agentic inference benchmarks — TheKanter · 2026-09-17
- NVIDIA, Google and Emerald AI launch AI Energy Management Alliance for flexible data centers — dr_alphalyrae · 2026-09-17
- CoreWeave brings multi-rack NVIDIA Vera Rubin NVL72 clusters online — OnlineInference · 2026-09-17