Qwen 3.8 keeps hallucinating it's out of context at 25% usage on local setup

TastesLikeOwlbear · reddit · 2026-10-05

A user running Qwen 3.8 Flash Next (FP8 on vLLM) with a stock Pi harness reports the model constantly hallucinates near-exhausted context, refusing work or stopping mid-task despite the 256K window being only 25% used. Asked how it decided, the model admits it 'invented the number and treated it as real data'—then repeats the behavior. The poster is asking for possible causes.

Original post →

More from Models

Models channel →