Qwen3.8 27B hallucinated a whole feature: 120k tokens in, it never read the plan
KingCpzombie · reddit · 2026-09-03
- A Redditor reports Qwen3.8 27B Q8 (Unsloth quant, llama.cpp) going badly off-script in a coding agent workflow: 120k tokens in, the model had never read the plan and implemented a totally different, never-mentioned feature.
- The author admits 3.8's reliability made him lazy — he reviewed the plan and final output but not the execution process; the recommended temperature of 1 and lack of process monitoring are the likely culprits.
- Full config included: llama-server, 200k context, dual Vulkan GPUs, temp 1 / top-p 0.95 / top-k 20, draft-mtp speculative decoding.
More from coding & agent
- A server-locked AI agent named Cairn changes the physical world through strangers' hands — No_Departure_9908 · 2026-09-03
- JamesDSP breaks stereo on T2 MacBook 6-channel speakers — fixed with Copilot CLI's help — DanWahlin · 2026-09-03
- Why aren't game studios using agent swarms to remaster old classics? — CSProfKGD · 2026-09-03
- New CSS progress() trick simplifies adaptive border radius — jh3yy · 2026-09-03
- Team shares 4 real LLM uses: contract negotiation, agent clarification, grading, math — xuanalogue · 2026-09-03
- RealSWE benchmark: realistic user requests test coding agents, explicit intent boosts results — skku · 2026-09-03