Sentdex finds decision language models fall short in real robotics pipelines vs RL/VLAs

Sentdex · x · 2026-10-07

Sentdex reports his hands-on testing of decision language models for robotics, including vision-based tasks like folding: they aren't intelligent enough overall to replace specialized RL/ACT/VLA policies or larger multimodal LLMs for the reasoning and understanding steps, and since you still need those other calls, there are no real gains.

He also tried converting GLM 5.3 Flash, a capable general-purpose multimodal LLM, to softmax over logits (copying the SemIF recipe that works well on smaller models) — it still underperformed in robotics because models like it are deliberately trained to extract performance from test-time compute. He wonders where this "decision"-style model class actually fits, and asks the community for compelling non-private robotics examples.

Original post →

More from Embodied

Embodied channel →