Running 30 concurrent Gemma 4 E4B calls locally: how to build it under $7k
Plane_Garbage · reddit · 2026-09-02
A Reddit user is looking for a local high-concurrency inference setup: roughly 30 simultaneous Gemma 4 E4B calls at 15k input / 3k output tokens each, ideally all finishing within 2 minutes, on a budget of $7,000 USD. Candidate options include multiple RTX 5060 Ti cards, a Mac Studio Max, DGX Spark, or several Intel B580s.
More from Infra
- antirez: DeepSeek v4 Flash is still the king of local inference, now with vision — antirez · 2026-09-02
- Cheapest hardware to run Qwen3.8 27B at Q8? Ascend 310 out of stock, MI50 questioned — Snoo-2768 · 2026-09-02
- Self-hosted LLM with SLERP-merged GRPO experts outperforms larger baseline, serving half of production traffic — t-tech · 2026-09-02
- AI Infrastructure Night event in San Francisco — glcst · 2026-09-02
- Google signs 396 MW geothermal deal to power AI amid energy crunch — VraserX · 2026-09-02
- Local Model Suitability MCP: Cuts Costs via Local Inference — modelcontextprotocol · 2026-09-02