DeepMind's Astra Robot Evaluated 98,000 Times Across 28 Simulated Environments

tomssilver · x · 2026-09-25

A new preprint accompanies the Astra robot videos: the authors built a large-scale sandboxed evaluation and ran 98,000 evaluations across 28 simulated environments, with a deliberately clean, strict experimental setup intended to serve future agent evaluation. Qualitative examples show agents devising surprisingly clever physical reasoning strategies.

Related event: 98,000 Simulated Trials Show Astra Robot's Surprising Physical Reasoning(3 posts)→

Original post →

More from Embodied

Embodied channel →