Qwen3.8-Flash-Next local quants hallucinate 'corrupted context' errors, reports Reddit user

arkham00 · reddit · 2026-09-03

A Reddit user reports that Qwen3.8-Flash-Next GGUF quantizations (Q4KM, Q3KXL, IQ4XS, with/without MTP, short and long contexts) run locally via llama.cpp on a Mac M2 Max 96GB frequently hallucinate "garbled text" — the model declares tool instructions or .md files corrupted, then runs extensive git/system checks even though the files are fine. The issue seems to be in the model's own context, with occasional stray Chinese characters and misspellings.

Notably, the model is self-aware about its errors and tries to recover; in one case it claimed a tool was corrupted, but when the user told it to just run the tool, it worked — the model apologized and continued. The author finds the base model strong on instruction-following, much better than predecessors, despite the preview-quality quirks.

Original post →

More from Models

Models channel →