Redditor spins up a 4x32GB V100 vLLM server, says local setup covers 90% of work

TrailFeatures · reddit · 2026-09-21

A Reddit user shares their local deployment: a Turnstone server with four 32GB V100s running 1Cat-vLLM, with Qwen3.8 Flash Next handling 90% of their workloads.

The project is forcing them to learn local LLM ops — a skill their company explicitly wants them to build so they can deploy it internally as well.

Original post →

More from Infra

Infra channel →