Qwen3.8-Flash-Next local quants hallucinate 'corrupted context' errors, reports Reddit user
arkham00 · reddit · 2026-09-03
A Reddit user reports that Qwen3.8-Flash-Next GGUF quantizations (Q4KM, Q3KXL, IQ4XS, with/without MTP, short and long contexts) run locally via llama.cpp on a Mac M2 Max 96GB frequently hallucinate "garbled text" — the model declares tool instructions or .md files corrupted, then runs extensive git/system checks even though the files are fine. The issue seems to be in the model's own context, with occasional stray Chinese characters and misspellings.
Notably, the model is self-aware about its errors and tries to recover; in one case it claimed a tool was corrupted, but when the user told it to just run the tool, it worked — the model apologized and continued. The author finds the base model strong on instruction-following, much better than predecessors, despite the preview-quality quirks.
More from Models
- Testing unreleased Gemini 3.8 Flash: no citations shown for top-of-funnel queries — gaganghotra_ · 2026-09-03
- Hidden-bug eval across 105 issues: Fable 5.1 finds 43, none fixes all — cost per model compared — PawelHuryn · 2026-09-03
- X open-sources new For You algorithm code: long dwell drives retrieval, bots can trigger account review — Kyrannio · 2026-09-03
- Gemini 3.8 Flash reverse-engineers Kerbal save files to build and land a Mun rocket — dosco · 2026-09-03
- User calls out model for double-standard answers on gendered scenario questions — Ribbitz_bow_tie27 · 2026-09-03
- Redditor Predicts Astra Model Release Tomorrow at 1pm PT Based on X Teasers — dolo937 · 2026-09-03