SegDAC: Boosting Visual RL Generalization with Dynamic Object Tokens
GlenBerseth · x · 2026-08-18
SegDAC is a novel approach in visual reinforcement learning that integrates pretrained segmentation models. Instead of operating on pixels, it generates object masks via text-grounded segmentation and extracts a variable-length set of dynamic object tokens. A Transformer-based actor-critic processes these tokens using segment positional encoding to preserve spatial information. Evaluated on 8 ManiSkill3 manipulation tasks, SegDAC improves visual generalization by 88% on the hardest perturbation settings while matching the sample efficiency of state-of-the-art methods. The project is fully open source.
More from Research
- RareBench results: Gemini 3.7 Flash leaps ahead, DeepSeek Pro shows no gain — danielmckinn0n · 2026-08-18
- Claude AI raises proven Riemann-zeta lower bound to 67.2% — thione · 2026-08-18
- Anthropic research documents failure patterns in frontier multiagent systems — thione · 2026-08-18
- Technical discussion on LLM vocab bottlenecks: smaller vocab, longer sequences? — LChoshen · 2026-08-18
- New research: are LLM agents time-aware? Can they estimate task duration? — maksym_andr · 2026-08-18
- Paper examines moderation in AI-generated sexual content communities — chaumian · 2026-08-18