GAVEL harness lifts Qwen3-8B from 41.2% to 91.8% on long-horizon robot tasks

dair_ai · x · 2026-09-20

dair-ai highlights the GAVEL paper: with no model changes, an external harness lifts Qwen3-8B from 41.2% to 91.8% on long-horizon robot tasks. GAVEL maintains an explicit graph world model of object relations, action preconditions/effects, and probabilistic beliefs about unobserved objects. Before executing an LLM-generated action, the graph predicts its outcome; violations are caught, fixes derivable from the world model are applied without re-querying the LLM, and only errors needing semantic reasoning go back to the model. On BEHAVIOR-1K (500 multi-task instructions), success rises from 19.9% to 92.6%, and reasoning over object-location distributions reorders subtasks, cutting travel distance 5.4%. The takeaway: much 'model weakness' is actually harness quality.

Original post →

More from Embodied

Embodied channel →