Dual RTX 5090 + 4070 TiS runs 27B FP8 with 258K context at 81 tok/s
Fz1zz · reddit · 2026-09-27
A Reddit user got Qwen3.8-27B block-FP8 (28.75 GiB) running on a 48GB setup combining an RTX 5090 (32GB) and RTX 4070 Ti Super (16GB) using vLLM 0.30.0 pipeline parallelism: the 4070 hosts layers 0-19 plus the vision encoder, while the 5090 handles layers 20-63, lmhead, and MTP.
Benchmarks:
- 1K in / 512 out, c1: 81.3 tok/s (ITL 36.7ms), 100 tok/s on code
- MTP acceptance 87%, full 262,144 context, fp8 KV
- 32K prefill: 3,160 tok/s; 145K fresh prefill: 68.6s; 258K prefill: 169s with 5090 peaking at 31,472 MiB
The 4070 Ti Super had been collecting dust for two months over PCIe x1 bandwidth concerns, but the tests showed the impact was negligible — useful data for anyone running large models on mixed GPUs.
More from Infra
- Three myths of hosted LLM inference: sticker prices, interchangeable endpoints, and self-hosting — TangeloOk9486 · 2026-09-27
- Random Attention: Salesforce and UIUC find random KV cache eviction rivals handcrafted signals — jiqizhixin · 2026-09-27
- Laptop engine streams a 35B model from SSD at 9.4 tok/s, beating GPT-OSS 20B — ImBadGuyInEveryStory · 2026-09-27
- World's fastest panel QR factorization on B200: how a GPU MODE contestant cracked chained dependencies — A_K_Nain · 2026-09-27
- TensorSharp open-source engine adds mixed document/image/video/audio inputs per request — fuzhongkai · 2026-09-27
- Pro-data center rally clashes with protesters as scholar defends AI infrastructure — neil_chilson · 2026-09-27