ExLlamaV3 3bpw local Flash benchmarks: 1500 tps prefill on a single RTX 5090
youcloudsofdoom · reddit · 2026-09-21
A Reddit user shares impressive results running Flash locally with ExLlamaV3/exl3 at 3bpw, saying it's replacing vllm/llama.cpp for them:
- 3x RTX 3090 + 128GB DDR4: 1500 tps prefill, 80 tps decode
- Single RTX 5090 + 128GB DDR4: 1500 tps prefill, 29 tps decode
- Both tested at 262k context with vision and speculative decoding enabled
- Quant quality at 3bpw looks good so far; a 4bpw comparison is planned
Recommended for anyone who was sleeping on exl3.
Related event: ExLlamaV3 3bpw Quantization Delivers Impressive Local Flash Performance(2 posts)→
More from Infra
- Cloudflare Quick Tunnels exposes localhost via one command, no account needed — Arindam_1729 · 2026-09-21
- AI Buildout Will Need Far More ABF Substrates, Says Analyst Ben Bajarin — BenBajarin · 2026-09-21
- Lumentum Shows Optics Linking Two AI Datacenters to Work as One Machine — BenBajarin · 2026-09-21
- WSJ: Powering the US AI build-out means walking power lines and sniffing for burning smells — xiaosun86 · 2026-09-21
- Running ComfyUI CPU-only: SDXL Turbo generates images in ~18s without GPU — tostane · 2026-09-21
- FAA to roll out AI tool for air traffic control starting Monday — Polymarket · 2026-09-21