Same open-weights model doom-loops in one engine, runs fine on llama.cpp

pand5461 · reddit · 2026-10-04

Testing a hyped inference engine, the author had an IQ3S quant answer a question about which open-weights models are least sycophantic while measuring long-generation throughput. The thinking trace degraded after 10k tokens into a doom loop of repeated fragments like "Need maybe maybe include..." and never produced a coherent answer.

The exact same model running in llama.cpp answered coherently with no loop — pointing the blame at the engine rather than the model. Takeaway: when an open model behaves strangely, "try llama.cpp" first.

Original post →

More from Infra

Infra channel →