MTPLX Ships New Version: Cache Copies Cut to One, Much Faster Decoding Past 140k Tokens
HankYeomans · x · 2026-10-04
Youssofal announced a new MTPLX release with major coding-focused improvements: legacy cache duplication reduced from 3 copies to 1 for better memory utilization, lower TTFT, sustained decode now near peak throughput, and much faster decoding past 140k tokens after fixing an issue where it drifted off the compiled route. Users with coding problems on V2.12 are encouraged to upgrade.
More from coding & agent
- TypeScript core member's first month at Cursor: 266 PRs, 15.5B tokens via agents — Vjeux · 2026-10-04
- Redditor vibe-codes a YouTube comment auto-poster with Gemini and local Qwen Coder in an hour — Intelligent_Many8573 · 2026-10-04
- Siqi Chen asks: what's the best production-validated agent memory architecture with temporality and dreaming? — msg · 2026-10-04
- SocialBu founder shows Claude MCP workflow for hands-off social media scheduling — usamaejazch · 2026-10-04
- Godot Port of CRT/VHS Glitch Effects Suite Ships With Agent-Friendly Docs — AIandDesign · 2026-10-04
- If LLMs write most code, why not design a programming language just for AI? — MyBeardHasThreeHairs · 2026-10-04