Custom llama.cpp fork pushes 27B models to 90+ TPS on an RTX 3090

Brief-Tap-6616 · reddit · 2026-09-15

A Reddit user released llamAmpere, a llama.cpp fork optimized for NVIDIA's Ampere architecture (RTX 30-series), with gains also carrying over to Blackwell and Lovelace cards.

Original post →

More from Infra

Infra channel →