$5,400 eBay 8x V100 Server Hits 200+ tok/s on a 27B Model with FlashAttention
MzCWzL · reddit · 2026-10-01
A Reddit user shared benchmarks running flash-next inference on a $5,400 eBay 8x V100 server:
- With TP=4 on just 4 GPUs, a 27B model reaches >200 tok/s (dflash) at 2.5k–3.5k tok/s prefill
- With image inputs enabled, KV cache supports roughly 120k context
- Runs Nvidia's nvfp4 checkpoints, unpacked to fp16 on the fly via a repo called 1Cat-vLLM, which the author heavily optimized
- Bring-up was done with the help of Claude Opus 5.5
More from coding & agent
- Git veteran warns: avoid SHA256 repos at Git 3 launch, they won't work for years — vmg · 2026-10-01
- Months of agent-driven dev: 'agents can't do big refactors' is outdated — dreamwieber · 2026-10-01
- Longtime Dev: "AI Can't Build Maintainable Architectures" Is a 2023 Take — dreamwieber · 2026-10-01
- Economist Paul Novosad: AI-edited code becomes unmanageable, so I wall it off — paulnovosad · 2026-10-01
- Local models can now power computer use agents, but regulated industries still lack a playbook — Ambitious_Fold_2874 · 2026-10-01
- Xiaomi MiMo-V2.6 report: one mixed GRPO run across all domains, $2.6M for Pro — SergioPaniego · 2026-10-01