Same open-weights model doom-loops in one engine, runs fine on llama.cpp
pand5461 · reddit · 2026-10-04
Testing a hyped inference engine, the author had an IQ3S quant answer a question about which open-weights models are least sycophantic while measuring long-generation throughput. The thinking trace degraded after 10k tokens into a doom loop of repeated fragments like "Need maybe maybe include..." and never produced a coherent answer.
The exact same model running in llama.cpp answered coherently with no loop — pointing the blame at the engine rather than the model. Takeaway: when an open model behaves strangely, "try llama.cpp" first.
More from Infra
- Coding agent costs down two months straight: LangChain CEO shares 3-step playbook — hwchase17 · 2026-10-04
- Rural data centers are set for a big US federal tax break — but some hyperscalers aren't biting — nordicinst · 2026-10-04
- Rural data centers are in for a big federal tax break, but hyperscalers seem lukewarm — Wired AI · 2026-10-04
- LocalMaxxing: community-run speed tests for local LLM rigs, 8.6K runs across 592 hardware setups — lxfater · 2026-10-04
- Lenovo's AI Express ships on-prem AI servers in 15 business days — shashib · 2026-10-04
- From 100KB to 2.5TB: a developer maps the full AI model spectrum and what each tier costs to run — abhishekkumar333 · 2026-10-04