Disable MTP when coding: Tests show speed drops drastically at long context
fbms2 · reddit · 2026-08-22
Tests indicate that Multi-Token Prediction (MTP) is counterproductive for coding. While MTP achieves 140t/s in chat mode on a 5090, coding speeds crash from 100t/s to 10-20t/s once the context exceeds 60k. Disabling MTP results in a stable 40t/s even at very high context lengths, despite a lower starting speed.
More from coding & agent
- Claude Agent Experiment Day 17: Self-Report on Memory Loss and Financial Autonomy — No_Departure_9908 · 2026-08-22
- 'Harness Engineering' Rises: Custom Scaffolds Become the Foundation of AI-Native Companies — omarsar0 · 2026-08-22
- Bootstrapped to $1M+ in 18 Months: A Look at 40 AI Agents Running the Business — aryanXmahajan · 2026-08-22
- Codex Builds Working Circuits Inside the Game 'Turing Complete' — Full CPU Next — Angaisb_ · 2026-08-22
- 2000 multimodal patent project rebuilt in a few Grok prompts 26 years later — Daniel_Farinax · 2026-08-22
- OpenAI lets MCP plugins ship bundled "skills" baked into ChatGPT and Codex — dfinke · 2026-08-22