95+ TPS and 262K context for Qwen 27B on a single RTX 3090 with LlamAmpere v0.4

Brief-Tap-6616 · reddit · 2026-09-29

Developer JakeATX released LlamAmpere v0.4, a Llama.cpp fork with Ampere-specific optimizations, hitting 95+ TPS and 262K context running Qwen3.8 27B (4.6bpw) on a single RTX 3090 — 10% faster than v0.3 with 10%+ more context. Closest rival vLLM stays within 10% but caps context lower. The post includes full build/run commands, notes EXL3 speedups (80% of XS-M quant speed), a Swift-qwen distill base with 1% performance loss, and benchmarks at temp=1 to avoid vanity temp=0 numbers. Open source on GitHub/HF.

Original post →

More from Infra

Infra channel →