Testing MTP combined with ngram-mod for coding speed

YetAnotherAnonymoose · reddit · 2026-08-18

A developer tested speculative decoding by combining MTP with ngram-mod in llama.cpp. The task involved repeating a code block 10 times.

Results:

Issue:

The author questions whether this is a parameter tuning issue or if the ngram combo is generally not worth it. This is particularly puzzling for models like Qwen3.8, which love to repeat 'thought' code blocks verbatim—an ideal case for ngram.

Original post →

More from Infra

Infra channel →