27B Model at 256k Context, 110+ tok/s on a Single RTX 5090 via focus-llama
Ok-Shower7286 · reddit · 2026-10-07
A developer shared focus-llama, a llama.cpp fork that runs Qwen3.8-27B (Q6KXL) on a single RTX 5090 with 256k logical context at a steady 110+ tokens/s.
- Vanilla llama.cpp allocates the full KV buffer up front, so VRAM scales with context and speed drops from 120 to 60 t/s as context grows
- focus-llama caps the physical KV buffer at 85k cells, offloading older chunks to an external focus-memory store fetched on demand
- Per-step attention cost is bounded by the buffer size rather than logical context length, keeping generation at 118–135 t/s
- It implements two Google DeepMind techniques (declarative attention and skill.state) plus adaptive MTP speculative decoding
- Context usage stays at 18–30%, eliminating full-day compaction pauses in coding harnesses like Cline or Qwen Code
The author posts the full configuration and logs — a valuable reference for long-context local coding workflows.
More from coding & agent
- OpenAI's Decisions API now takes image input — one creator picks YT thumbnails for $0.13 — stevenheidel · 2026-10-07
- Embedder agent migrates nRF52 firmware to nRF54/Zephyr in under 1 hour with a single prompt — ycombinator · 2026-10-07
- Brain Co launches Synapse, a design system that encodes judgment for AI agents — LVidegaray · 2026-10-07
- Vibe coding veterans now at a disadvantage: just tell the model what you want — TheMoonMidas · 2026-10-07
- Lerna Plugin Updated: Route GitHub Copilot CLI Calls to Microsoft Foundry — unixterminal · 2026-10-07
- Agents are in their bundling phase; specialized unbundled agents may be next — signulll · 2026-10-07