LlamAmpere release: run 27B models with 200K context on 12GB cards, new KVaRN cuts KLD ~40%
Brief-Tap-6616 · reddit · 2026-10-10
LlamAmpere ships a new open-source release targeting 12GB Ampere cards, plus 3-4% speedups and smaller runtime for 3090 users (340K ctx with YaRN).
- New Staged + Journaled KVaRN variant reduces KLD vs. paper-faithful and competitive versions by 40%.
- 4/4 is the new recommended default for larger cards; 3/3 (0.001 nats KLD) and 3/2 (0.0024 nats) quant variants support 205K-230K ctx at 11GB for maximum context.
- Tested model: a 2.3bpw fusion of swift-1.5-uncensored and mirai's 2.5bpw model, retaining 85% of BF16 performance on LiveCode Bench across 7 runs, 65-70 tps on 3080/ti thanks to MTP.
- Author reports it's genuinely usable for standard agentic tasks, subjectively between BF16 Qwen3.5 and 3.6. MIT-licensed, models on Hugging Face.
More from Infra
- Agent-built system hits 2,242 tok/s on AMD MI300As, 2.33× faster than SGLang in 105 hours — bariskasikci · 2026-10-10
- Bespoke serving systems becoming the only viable path as hardware-model combos explode — bariskasikci · 2026-10-10
- Napkin math on the AI bubble: 100x cheaper models still mean 1000x more GPUs — gabriel1 · 2026-10-10
- Super Micro contractor pleads guilty in $2.5B scheme smuggling Nvidia AI chips into China — Polymarket · 2026-10-10
- Cloudflare acquires Deno, will maintain runtime for only one more year — Simon Willison · 2026-10-10
- How apps scale: 2006 bigger servers, 2016 clusters, 2026 rewrite in Rust — tristanbob · 2026-10-10