Running a 27B model locally on 2x RTX 5090 with vLLM
piddlefaffle12 · reddit · 2026-09-25
A Reddit user shared a screenshot of running a 27B-class open model locally via vLLM on two RTX 5090 GPUs under CachyOS — a practical consumer dual-GPU local inference setup reference.
More from Infra
- Smartphones eat ~30% of global DRAM and NAND supply — the fix? Stop yearly phone releases — AlpinDale · 2026-09-25
- TileRT and AMD hit 469 tok/s decode on GLM-5.3 with vLLM on 8x MI355X, 40% faster than GB300 — vllm_project · 2026-09-25
- AI now beats humans at some TPU design tasks, but is still seen as just a tool — burny_tech · 2026-09-25
- More Budget 4-GPU Inference Tricks: x8 Splitters and m.2-to-x4 Adapters — TheZachMueller · 2026-09-25
- CLion 2026.2.3 adds NVIDIA CUDA Tile C++ support with dedicated inspections — blelbach · 2026-09-25
- AMD claims edge box runs 2.3x more workloads than Jetson T5000-class hardware — shashib · 2026-09-25