Qwythos-9B-v2 Fixes Repetition Loop Issue
techlatest_net · reddit · 2026-07-14
The Qwythos-9B-v2 GGUF card updates model usage instructions and improvements: the new version trains away the loop degradation issue under greedy or low-temperature decoding from the previous version, reducing loop behavior from 6.7% to 0%, while restoring the native MTP head and cleaning up the identity prompt; knowledge and reasoning capabilities remain at least as good as the base version.
The training fix method is called FTPO (Final-Token Preference Optimization), which focuses preference optimization only on the starting token that triggers the repetition loop, guiding the model toward coherent alternatives while minimally affecting the rest of the distribution. The post also provides loading instructions for llama.cpp, Ollama, LM Studio/jan/KoboldCpp, as well as a list of required files for normal text weights, MTP weights, and the vision projector for image input.
More from Models
- NVIDIA says Nemotron 3 Ultra scored 30/42 on the 2026 IMO problems — NVIDIAAI · 2026-07-22
- OpenAI is reportedly briefing U.S. lawmakers on its next model family — kimmonismus · 2026-07-22
- Muse Spark 1.1 lands at 1495 on Text Arena with standout agentic-coding price performance — ycombinator · 2026-07-22
- Advanced AI Models Are Becoming Impossible to Plug and Play — emollick · 2026-07-22
- Google Gemini's AI Problem: No Leading Model for Core Workloads — bindureddy · 2026-07-22
- Model Offers 1M Token Context Window at Just $0.33/1M Tokens — MickeySteamboat · 2026-07-22