Running 30 concurrent Gemma 4 E4B calls locally: how to build it under $7k

Plane_Garbage · reddit · 2026-09-02

A Reddit user is looking for a local high-concurrency inference setup: roughly 30 simultaneous Gemma 4 E4B calls at 15k input / 3k output tokens each, ideally all finishing within 2 minutes, on a budget of $7,000 USD. Candidate options include multiple RTX 5060 Ti cards, a Mac Studio Max, DGX Spark, or several Intel B580s.

Original post →

More from Infra

Infra channel →