Video understanding models still fail on mirror reflections and contradictions

Kyrannio · x · 2026-07-21

Video understanding models still hallucinate on mirror scenes

In reply to a thread about video understanding, the author says the field still badly needs better models. Their example is simple: generate an AI video with a mirror reflection, feed it into an LLM such as GPT-5.6, Gemini, or Fable, and ask whether the scene makes sense.

The result, they say, is that the models produce weird hallucinations and contradictory reasoning. The post is not about a fix, but about the gap between current video understanding claims and actual reasoning reliability in tricky visual setups.

Related event: Video Understanding Models Still Struggle with Mirror Reflections(2 posts)→

Original post →

More from Models

Models channel →