50+ tps on RTX 5060 Ti via MTP fails to reproduce; Blackwell llama fork hits 20-35 t/s
Arany8 · reddit · 2026-09-28
- A Reddit user tried and failed to reproduce a claimed 50+ tokens/s setup on an RTX 5060 Ti using MTP speculative decoding, questioning whether the original claim is real.
- Along the way they released beellama.cpp, a freshly built llama.cpp fork targeting Blackwell (sm120), with a full llama-server config: kvarn3 KV cache + flash-attn, 98304 context, draft-mTP (draft-n-max 4, ubatch 128), and q80/q40 KV quantization.
- Their real-world result is 20-35 t/s, well short of the claimed 50+; the thread is still working out the right settings.
More from Infra
- MLX MoE Layer Gets 1.5x Faster via Better Tile Scheduling in Grouped Matmul — awnihannun · 2026-09-28
- Cloudflare incident: skipped block zeroing leaked tenant data across 18 of 24 containers — arpit_bhayani · 2026-09-28
- On RTX 5090, Qwen 27B hits 200 TPS but Flash next only 50: what model sits between for coding? — MasterNomie · 2026-09-28
- Lumen Launches On-Demand Dedicated Internet Up to 100 Gbps at 10M US Sites — shashib · 2026-09-28
- Local AI comes in two flavors: laptop-scale for the masses vs SMB on-prem setups — TheZachMueller · 2026-09-28
- Cloudflare's agent-first Kitesurf browser adds WebMCP, passes 730k WPT subtests, runs in terminals — Cloudflare Blog · 2026-09-28