Mirage Hallucinations Linearly Decodable from Model Activations
ziv_ravid · x · 2026-07-04
Research shows that a model's Mirage behavior (fabricating images/details) can be linearly decoded from its internal activations, even when the images actually exist. Pure text baselines fail to recover this signal. The study also suggests that activation signals can help distinguish between different types of Mirage behaviors.
More from Research
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11