Dev launches Lorivo: one GPU server serves many LoRA adapters via vLLM
TheOneWhoWil · reddit · 2026-09-13
A Redditor launched Lorivo, a serverless hosting platform for LoRA adapters built on vLLM. Since a rank 8–32 LoRA shares 99%+ weights with its base model, Lorivo routes adapters to one shared GPU server per base model, leveraging vLLM batching and fused LoRA kernels, and exposes OpenAI-compatible endpoints. Deploy takes two CLI commands. Currently hosts Qwen 3.5 4B free (32k context) on the author's own GPU, with $1,000 in AWS credits to add more models; no chats or requests are stored, only token counts and timestamps.
More from Infra
- DeepMind Chief Strategist: AI Infrastructure Spending Is a Bet on Recursive Self-Improvement — rohanpaul_ai · 2026-09-13
- AI valuations can't all be right: memory at 3-5x PE vs premium infrastructure, says Gavin Baker — rohanpaul_ai · 2026-09-13
- Draw Things launches Local Code beta: local coding agents on Mac at ~980 tok/s prefill — liuliu · 2026-09-13
- Macrocosmos launches IOTA for liquid training on scattered, disaggregated compute — markjeffrey · 2026-09-13
- "The CUDA moat is gone": Japanese neocloud ai& deploys Tenstorrent at scale — DavidBennett__ · 2026-09-13
- Patched vLLM+FlashInfer Pushes Gemma 4 31B to 150 tok/s on a Single B300, Beating SGLang — abhijithneil · 2026-09-13