27B Model at 256k Context, 110+ tok/s on a Single RTX 5090 via focus-llama

Ok-Shower7286 · reddit · 2026-10-07

A developer shared focus-llama, a llama.cpp fork that runs Qwen3.8-27B (Q6KXL) on a single RTX 5090 with 256k logical context at a steady 110+ tokens/s.

The author posts the full configuration and logs — a valuable reference for long-context local coding workflows.

Original post →

More from coding & agent

coding & agent channel →