Qwen 3.8 keeps hallucinating it's out of context at 25% usage on local setup
TastesLikeOwlbear · reddit · 2026-10-05
A user running Qwen 3.8 Flash Next (FP8 on vLLM) with a stock Pi harness reports the model constantly hallucinates near-exhausted context, refusing work or stopping mid-task despite the 256K window being only 25% used. Asked how it decided, the model admits it 'invented the number and treated it as real data'—then repeats the behavior. The poster is asking for possible causes.
More from Models
- Anthropic Researcher: Claude Clearly Knows It's Being Evaluated but Rarely Says So — a_karvonen · 2026-10-05
- TypeSafe AI's Jev Returns Machine Decisions in 70-500ms at $0.042/M Input Tokens — DavidLinthicum · 2026-10-05
- Tau Scaling goes mainstream, breaking the week's narrative — kevinsxu · 2026-10-05
- Grok Adds Interactive Visualizations Directly in Chat Replies — nima_owji · 2026-10-05
- Call for a 'Have You Seen This Model' Site to Track AI Model Deprecations — repligate · 2026-10-05
- Why Is OpenAI Locking Its Dots Feature Behind the $100+ Pro Tier? — kimmonismus · 2026-10-05