Mirage Hallucinations Linearly Decodable from Model Activations

ziv_ravid · x · 2026-07-04

Research shows that a model's Mirage behavior (fabricating images/details) can be linearly decoded from its internal activations, even when the images actually exist. Pure text baselines fail to recover this signal. The study also suggests that activation signals can help distinguish between different types of Mirage behaviors.

Original post →

More from Research

Research channel →