Fixing MTP head cuts Ornith1.5 35B inference time by 33%
frankentriple · reddit · 2026-08-22
The author spliced a trained MTP (Multi-Token Prediction) head onto an APEX requantized version of Ornith1.5 35B, addressing the untrained head issue in the original release. While the tokens per second (TPS) only increased slightly from 60 to 64 (+6.7%), the end-to-end task completion time dropped by 33% (from 21s to 14s), significantly boosting efficiency. The optimized model is available on Ollama, with testing methodology and detailed results published on GitHub.
More from Infra
- New site tracks data center impact on electricity prices: Effects are minimal — soumitrashukla9 · 2026-08-23
- Enterprise LLM Gateway 2026: Support for Reasoning Models — llogicnotfound · 2026-08-23
- Expanding Inference Server: AMD/Intel vs NVIDIA CUDA Stack — AdSafe4047 · 2026-08-23
- The Real AI Risk Isn't the Model: Permissions, Inputs, and Presentation — bigdata · 2026-08-23
- a16z: agents burn nearly 5x the tokens humans do, up 14x since February — a16z · 2026-08-23
- SiliconScope: Sudoless monitor tracks Apple Silicon ANE, Media Engine, and memory bandwidth — tom_doerr · 2026-08-23