SEAR benchmark tests how LLM agents self-evolve on robots, from DeepSeek to Astra
xiye_nlp · x · 2026-10-08
Researchers introduced SEAR, a benchmark and framework for studying how LLM agents self-evolve on robots. The loop: within a time budget, an agent and its skill library update a policy, run it in simulation or on hardware, receive judge feedback, and carry lessons into the next task. It studies three modes at once: code-as-policy, agent-as-policy, and autoresearch (training its own VLA). Findings: not just Astra but even text-only LLMs like DeepSeek can control robots; Astra excels at live watching and correction, Fable at building tools and policies in code first. Over 4 years of cumulative agent evolution time was invested.
More from Embodied
- Comma_ai user praises driving data setup, but car chewed through the cables — fforres · 2026-10-08
- ~200 production Cybercabs without steering wheels spotted staging in Houston — JOBhakdi · 2026-10-08
- LEDGER builds persistent 3D object memory from egocentric video, answers questions without rewatching — mangahomanga · 2026-10-08
- NavSafe-∞: photorealistic closed-loop benchmark exposes safety gap across 20 E2E driving policies — zhoubolei · 2026-10-08
- Polymarket bets open on whether Meta ships its Tamagotchi-style Muse Charm AI device by December — Polymarket · 2026-10-08
- Microsoft's $5,999 Surface RTX Spark Dev Box preorders open, ships November — tomwarren · 2026-10-08