Humans and VLMs produce surprisingly similar image reconstructions from memory

ducha_aiki · x · 2026-09-09

Researcher duchaaiki shares an observation: when humans and VLMs reconstruct an image from memory or a text description, their outputs are strikingly similar. The thread also cites Jakob Engel's remarks on world models, suggesting the finding bears on whether VLMs build human-like internal representations.

Related event: DeepMind's Jakob Engel on World Models: Egocentric Multimodal Data Is the Future(4 posts)→

Original post →

More from Models

Models channel →