Building a Local Inference Server with $100 GPUs
Boricua-vet · reddit · 2026-07-12
The post shows a configuration to piece together local inference capability with $100-level GPUs, targeting 20GB VRAM, 448GB/s bandwidth, and support for 3 concurrent users.
Logs show the author loading Qwen3.6-35B-A3B-UD-IQ4XS.gguf on llama.cpp / llamaserver, encountering:
- ngpulayers too high causing parameters to not fit VRAM
- Context length nctxseq 32768 < nctxtrain 262144, cannot fully utilize training context
- Prompt cache, context checkpoints, speculative decoding and other server capabilities enabled
Overall, it discusses how to make a usable local model service with low-cost hardware, exposing real-world deployment issues with memory, context, and cache configuration.
More from Infra
- PyTorch Day Korea 2026 launches first offline conf, CFP closes Sept 13 — PyTorch · 2026-09-11
- Local LLM server dilemma: 4x CMP-170HX (price up 53% in 20 days) vs Mac Studio M5 Ultra — rumboll · 2026-09-11
- llama.cpp lands Flash Attention tuning for RDNA4, big prefill gains on AMD — pmttyji · 2026-09-11
- Your p99 latency benchmark may be lying: a deep dive into coordinated omission — Franc0Fernand0 · 2026-09-11
- Running MiniMax H3 on 12GB VRAM: quantization, Turbo LoRAs and attention backends compared — Possible_Mood676 · 2026-09-11
- Spomin: live KV cache compaction squeezes 500k tokens of context into 180k resident — wgaca2 · 2026-09-11