He built a €18k server with 768GB VRAM—now 2-trillion-parameter open models may leave it behind
myreala · reddit · 2026-09-04
A Reddit user spent €18k building an EPYC server with twelve 64GB cards (768GB VRAM total) plus 256GB RAM to run frontier open-source models locally. Now he fears he's already obsolete: next-gen open frontier models look likely to hit 2 trillion parameters or more.
In his view only GLM 5.3 still fits, and it will lag once Astra ships; GLM 6 could be two to three times larger, while Qwen-max, Kimi, and DeepSeek V4 Pro are already too big to consider.
He's weighing whether to sell the excess GPUs and settle for flash models on fewer cards. He had hoped to build a business on it, but admits the same business could probably run on a flash model with more engineering effort—a textbook dilemma for hobbyist-scale compute in an era of rapidly ballooning models.
More from Infra
- Nvidia A100 pre-training pipeline goes live on a 24/7 public livestream — wavefnx · 2026-09-04
- ADSP Episode 302: Mark Saroufim on PyTorch, GPU MODE, and automating AI research — blelbach · 2026-09-04
- Spotify's Portal Cut Claude Code Token Usage by 90% With a Two-Mode Router — rseroter · 2026-09-04
- Vyact: Open-Source Desktop Workspace Unifying Local LLMs, RAG, and Browser Context — vyact · 2026-09-04
- At what context depth does KV quantization start to hurt? An F16 vs Q8/Q4 parity PoC — Slight_Analysis_5414 · 2026-09-04
- KV caching: the fundamental optimization behind autoregressive LLM inference — alec_helbling · 2026-09-04