SEAR benchmark tests how LLM agents self-evolve on robots, from DeepSeek to Astra

xiye_nlp · x · 2026-10-08

Researchers introduced SEAR, a benchmark and framework for studying how LLM agents self-evolve on robots. The loop: within a time budget, an agent and its skill library update a policy, run it in simulation or on hardware, receive judge feedback, and carry lessons into the next task. It studies three modes at once: code-as-policy, agent-as-policy, and autoresearch (training its own VLA). Findings: not just Astra but even text-only LLMs like DeepSeek can control robots; Astra excels at live watching and correction, Fable at building tools and policies in code first. Over 4 years of cumulative agent evolution time was invested.

Original post →

More from Embodied

Embodied channel →