Qwopus 3.8 27B Flash fine-tune ships: 12.8% faster decoding, 80.7% MTP acceptance on Qwen3.8-27B
EAccelerate_42 · x · 2026-09-04
Community fine-tune Qwopus 3.8 27B Flash (Apache-2.0, GGUF on Hugging Face), built on Qwen3.8-27B, targets cheaper and faster long-running agent workloads.
- 12.8% faster decoding, 80.7% MTP acceptance rate; CoT is far more concise (sometimes 10x+ shorter) and the model chooses when to reason more or less
- Runs well with uncapped thinking in agentic workflows like Claude Code; 37 tps on an old V100, 100 tps on a 5090 with MTP
- Supports vision, tool use and code generation; aimed at local inference and wall-clock efficiency
More from Models
- Deep Learning Weekly #471: Claude Fable 5.1 launch, production-parity LLM evals, alignment paper — dl_weekly · 2026-09-05
- OpenAI confirms Astra counts toward normal plan usage, users can allocate 100% of quota — soumitrashukla9 · 2026-09-05
- RareBench eval: Claude Fable 5.1 tops rare-disease diagnosis while Nemotron 3 Ultra scores 0% — danielmckinn0n · 2026-09-05
- GPT-6 Astra Beats 5.6 Sol Pro (Max) on FrontierMath T4; Open Models Seen 18 Months Behind — inductionheads · 2026-09-05
- TheZvi breaks down the Claude Fable 5.1 system card: 200+ pages of safety evals — TheZvi · 2026-09-05
- COLM paper: legible chain-of-thought steps aren't necessarily important — LauraRuis · 2026-09-05