Benchmarking llama.cpp on 5090+3090: ubatch, lazy-mode and MTP tuning for 60k-260k context

Blindax · reddit · 2026-09-22

A Reddit user benchmarked llama.cpp running Qwen3.8-27B and Flash-Next at 60k-260k context on a Ryzen 9800X3D, 128GB RAM, dual RTX 5090+3090 setup, with reusable tuning findings:

Original post →

More from Infra

Infra channel →