Grounding LLMs with JEPA world models trained in simulation: a research proposal

Full_Promotion4522 · reddit · 2026-09-03

A Redditor proposes a research idea: LLMs learn statistical token relations ("falls"↔"gravity") without grounded physical intuition — essentially the Mary's Room problem. The proposal:

The hypothesis: downstream learning speeds up significantly since the LLM needn't rediscover that objects fall. Adjacent work: V-JEPA (predicts future frame representations for video), DreamerV3 (latent world models for RL), but this exact combination appears unexplored.

Open questions posed to the community: any missed prior work? What's the right interface (prompt concatenation vs cross-attention)? Will the sim-to-real gap kill transfer? Worth building a small prototype?

Original post →

More from AGI Musings

AGI Musings channel →