Reward-DAgger: generalist reward models enable task-agnostic runtime monitoring for robots

ebiyik_ · x · 2026-10-07

The author publishes his promised paper, Reward-DAgger: using generalist reward/progress models (like Robometer) for runtime monitoring in robot deployment.

The problem: existing methods rely on task-specific failure models or policy-dependent uncertainty, requiring retraining or recalibration whenever tasks or policies change. Reward-DAgger asks whether a gate can transfer across both.

The takeaway: generalist reward models are now good enough to replace task-specific heuristics for deciding when to request corrections.

Original post →

More from Embodied

Embodied channel →