SegDAC: Boosting Visual RL Generalization with Dynamic Object Tokens

GlenBerseth · x · 2026-08-18

SegDAC is a novel approach in visual reinforcement learning that integrates pretrained segmentation models. Instead of operating on pixels, it generates object masks via text-grounded segmentation and extracts a variable-length set of dynamic object tokens. A Transformer-based actor-critic processes these tokens using segment positional encoding to preserve spatial information. Evaluated on 8 ManiSkill3 manipulation tasks, SegDAC improves visual generalization by 88% on the hardest perturbation settings while matching the sample efficiency of state-of-the-art methods. The project is fully open source.

Original post →

More from Research

Research channel →