Custom llama.cpp fork pushes 27B models to 90+ TPS on an RTX 3090
Brief-Tap-6616 · reddit · 2026-09-15
A Reddit user released llamAmpere, a llama.cpp fork optimized for NVIDIA's Ampere architecture (RTX 30-series), with gains also carrying over to Blackwell and Lovelace cards.
- Recommended config hits 90+ TPS up to 100K tokens (at temp 1), with context support up to 240K
- Ships with a Qwen3.8-27B IQ4XS GGUF quant: roughly 4KM quality but considerably faster
- Author claims 80% speedup over near competitors at 200K context, making local 27B inference faster than API speeds on a single 3090
More from Infra
- US's #2 law firm Latham & Watkins is building its own in-house Nvidia AI stack — MikeBirdTech · 2026-09-15
- Sunk Cost: A calculator for when a local LLM rig pays for itself vs renting tokens — rlindsey123 · 2026-09-15
- GLM 5.3 Flash EXL3 adds TP=3 mode: 28% faster decode across three DGX Sparks — HankYeomans · 2026-09-15
- Why the AI buildout will continue regardless of Dario's pacing call — AccBalanced · 2026-09-15
- NVIDIA tops TIME's World's Best Companies of 2026 list for second straight year — DeryaTR_ · 2026-09-15
- Researcher open-sources theseus, a human-language architecture research framework — pratyusha_PS · 2026-09-15