Optimal Settings for Llama.cpp + Qwen 3.8: n-max 4 Fastest
GodComplecs · reddit · 2026-08-31
Testing on RTX 3090 with the Qwen 3.8 UD Q4 KM model revealed that n-max 4 combined with spec draft min p 0.7 yields the fastest speed for harness workflows. The configuration supports 205k context with Flash Attention enabled. Benchmarks showed a generation speed of 70 tks in testing, outperforming previous mtp2 settings.
More from coding & agent
- GitHub Hit: Security Skill Router for AI Agents with Self-Evolving KB — udmrzn · 2026-08-31
- How Mole keeps 110k lines of AI-generated Swift from rotting: 6 engineering rules — dotey · 2026-08-31
- Next Big Languages Will Prioritize LLMs Over Humans — jfischoff · 2026-08-31
- Comparing Multi-Agent Coding Workspaces: Tutti vs. Paseo for Teams — Careless_Cress_865 · 2026-08-31
- New Claude Code Version Launches 20% Faster — BLUECOW009 · 2026-08-31
- LFS pause breaks AI training; configurable progress bar coming soon — janusch_patas · 2026-08-31