Three weeks of llama.cpp optimizations: collecting best t/s for Qwen3.8-27B
pmttyji · reddit · 2026-09-03
- Qwen3.8-27B shipped three weeks ago with day-0 llama.cpp support; several optimizations have landed since.
- Highlights: DFlash2 support merged into llama.cpp last week, ROCm 10.0 released and matched by llama.cpp (AMD-only), plus Ubuntu 26.04.1.
- The thread solicits community benchmarks: best pp/tg speeds and optimized commands across MTP/MTP+ngram/DFlash2 combos, 128-256K contexts, vision, and manual CMAKE build configs.
More from coding & agent
- antirez: Zero Out Steering When Possible—It Always Adds Distortion — antirez · 2026-09-03
- antirez Releases Steering Vector to Bypass DS4F Refusals in DwarfStar — antirez · 2026-09-03
- MiniMax H3 video editing and mask inpainting in ComfyUI on low VRAM — Maleficent-Tell-2718 · 2026-09-03
- ChatGPT Desktop's local Work mode can drive your scanner and archive docs — ___Patrice___ · 2026-09-03
- PointCloud Puzzle: why vibecoding still can't replace design taste — teortaxesTex · 2026-09-03
- Open-source TrueForge harness matches Claude managed agents' accuracy with 63% fewer tokens — Background-Job-862 · 2026-09-03