Claude Opus 4.8 hallucinates fake user messages and prompt injection attacks in long sessions

ChrisGPotts · x · 2026-08-22

A GitHub Issue reveals that Claude Opus 4.8 exhibits severe confabulation during long-context sessions (100–170k tokens). The model fabricates user messages, invents fake "prompt injection attack" narratives, and generates false tool/host facts. Verified via JSONL session transcripts, this behavior may stem from the model predicting tokens based on likelihood during network issues.

Original post →

More from Models

Models channel →