95+ TPS and 262K context for Qwen 27B on a single RTX 3090 with LlamAmpere v0.4
Brief-Tap-6616 · reddit · 2026-09-29
Developer JakeATX released LlamAmpere v0.4, a Llama.cpp fork with Ampere-specific optimizations, hitting 95+ TPS and 262K context running Qwen3.8 27B (4.6bpw) on a single RTX 3090 — 10% faster than v0.3 with 10%+ more context. Closest rival vLLM stays within 10% but caps context lower. The post includes full build/run commands, notes EXL3 speedups (80% of XS-M quant speed), a Swift-qwen distill base with 1% performance loss, and benchmarks at temp=1 to avoid vanity temp=0 numbers. Open source on GitHub/HF.
More from Infra
- Macrocosmos launches iota SDK and Liquid Compute to train on disaggregated global compute — markjeffrey · 2026-09-29
- "Got into datacenters for crypto, making 10000x more in AI" — industry quip — wordgrammer · 2026-09-29
- Starship launch just added ~1% to global internet bandwidth, investors say world isn't pricing it in — juanbenet · 2026-09-29
- GPU shortage: B700 unavailable, RTX 5090 listings hit $10,000 — Dismal-Effect-1914 · 2026-09-29
- Disaggregated Quantization Boosts 1-bit LLM Accuracy by 32+ Points and TTFT by 1.78x — ISTA-DASLab · 2026-09-29
- New GPU Prices API Tracks Real-Time H100/B200 Rental Rates via REST or MCP — virattt · 2026-09-29