El Pulpo 0.1.0: A 12.7MB Proxy and Load Balancer for Local LLM Inference

zaytzev · reddit · 2026-09-16

A developer released El Pulpo 0.1.0, an open-source proxy/load balancer for local LLM inference, in a 12.7MB image. It unifies model provider settings across agents (no more reconfiguring when switching models or networks like Tailscale), adds token usage monitoring to see if devices pay for themselves, and load-balances across instances — built around the author's Qwen3.8 Flash Next setup.

Original post →

More from Infra

Infra channel →