GAVEL harness lifts Qwen3-8B from 41.2% to 91.8% on long-horizon robot tasks
dair_ai · x · 2026-09-20
dair-ai highlights the GAVEL paper: with no model changes, an external harness lifts Qwen3-8B from 41.2% to 91.8% on long-horizon robot tasks. GAVEL maintains an explicit graph world model of object relations, action preconditions/effects, and probabilistic beliefs about unobserved objects. Before executing an LLM-generated action, the graph predicts its outcome; violations are caught, fixes derivable from the world model are applied without re-querying the LLM, and only errors needing semantic reasoning go back to the model. On BEHAVIOR-1K (500 multi-task instructions), success rises from 19.9% to 92.6%, and reasoning over object-location distributions reorders subtasks, cutting travel distance 5.4%. The takeaway: much 'model weakness' is actually harness quality.
More from Embodied
- Scobleizer: Tesla FSD drives Austin with no humans onboard, AI racers already beat humans on F1 tracks — Scobleizer · 2026-09-20
- Musk confirms every Starlink V3 satellite will carry an Nvidia Vera Rubin NVL72 — 100k sats, 25GW — ns123abc · 2026-09-20
- Odyssey-3: One Pretrained World Model Adapts to Different Robot Arms With Hours of Demos — ChongZzZhang · 2026-09-20
- iPhone 18 Pro motherboard weighs just 13.1g yet packs PC-class compute — SumitGup · 2026-09-20
- How did Apple Silicon get 50% faster in three years? – Daniel Lemire — ibobev · 2026-09-20
- Snap teams with Salesforce, Amazon and Nvidia to push Specs glasses into the enterprise — matt_slotnick · 2026-09-20