Single RTX 5090 Benchmark: 27B Model at 616 Tok/s with 262K Context
EAccelerate_42 · x · 2026-08-26
- Configuration: Running Qwen2.5-27B on a single RTX 5090 (32GB, Blackwell).
- Performance: Achieves 616 tok/s aggregate throughput at 262K context length with 4-way concurrency, using NVFP4 weights, KV Cache, and DFlash2 (K=7) speculative decoding.
- Technical Details: Provides patches for vLLM v0.27.1, Dockerfile, and build scripts. Supports tool calling and an optional CPU vision sidecar.
- Deployment: Includes a two-command deploy process and comprehensive benchmark evidence.
More from Infra
- Spain plans stricter rules for data centers on water, energy, and security — Polymarket · 2026-08-26
- TorchMorph: CUDA-Accelerated Morphological Transforms for PyTorch — kornia_foss · 2026-08-26
- Shopify CEO open-sources walgit, a database-free Git server using S3 as the repo — Shruti_0810 · 2026-08-26
- Jalapeno chip shows strength, revealing Nvidia's inference architecture weaknesses — beffjezos · 2026-08-26
- Australia's PM backs down on requiring AI datacentres to run fully on renewable energy — nordicinst · 2026-08-26
- Running Qwen3.8-27B for local coding on 16GB VRAM: full setup guide — Due-Project-7507 · 2026-08-26