Fixing MTP head cuts Ornith1.5 35B inference time by 33%

frankentriple · reddit · 2026-08-22

The author spliced a trained MTP (Multi-Token Prediction) head onto an APEX requantized version of Ornith1.5 35B, addressing the untrained head issue in the original release. While the tokens per second (TPS) only increased slightly from 60 to 64 (+6.7%), the end-to-end task completion time dropped by 33% (from 21s to 14s), significantly boosting efficiency. The optimized model is available on Ollama, with testing methodology and detailed results published on GitHub.

Original post →

More from Infra

Infra channel →