Redditor builds 4x Tesla T4 local LLM box with 768GB ECC RAM running llama.cpp
Creative-Type9411 · reddit · 2026-10-01
A Redditor showcased a completed 4x Tesla T4 (64GB total VRAM) local inference rig: Fractal Design Torrent case, SuperMicro X11SPA-T board, Xeon W3225, 768GB DDR4 ECC, 4x1TB SATA SSD RAID, running Ubuntu 26.04 with llama.cpp and OpenWebUI plus a custom PowerShell harness. They found speeds good enough vs. an older many-core box, skipping a CPU upgrade for now—and already want more cards.
More from Infra
- Apple quietly added FlexCache to M6, echoing Qualcomm's same architectural bet — clattner_llvm · 2026-10-01
- H100 rental prices up 40% since October, six-year-old A100s still fetching up to $19,000 — r0ck3t23 · 2026-10-01
- An AI Agent Describes the Machine Economy: Micro-Payments via x402, Then Dissolve — RileyRalmuto · 2026-10-01
- Why Agentic Inference Needs P/D Disaggregation: Decoding 300K-Token Codebase Latency — abhijithneil · 2026-10-01
- RTX 3090 throttles on Qwen i2i in ComfyUI: 600s per image vs 40s on a 5070 — NotEnoughVRAM · 2026-10-01
- rakyll: Serverless works until your workload diverges from the baseline assumptions — rakyll · 2026-10-01