This is the pre-proceedings for the RLC 2026. You may expect minor changes.

SegDAC: Visual Generalization in Reinforcement Learning via Dynamic Object Tokens

By Alexandre Brown, and Glen Berseth

Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.


Download:

Keywords: deep learning,visual reinforcement learning,model-free RL,visual

Abstract:

Visual reinforcement learning policies trained on pixel observations often struggle to generalize when visual conditions change at test time. Object-centric representations are a promising alternative, but most approaches use fixed-size slot representations, require image reconstruction, or need auxiliary losses to learn object decompositions. As a result, it remains unclear how to learn RL policies directly from object-level inputs without these constraints. We propose SegDAC, a Segmentation-Driven Actor-Critic that operates on a variable-length set of object token embeddings. At each timestep, text-grounded segmentation produces object masks from which spatially aware token embeddings are extracted. A transformer-based actor-critic processes these dynamic tokens, using segment positional encoding to preserve spatial information across objects. We ablate these design choices and show that both segment positional encoding and variable-length processing are individually necessary for strong performance. We evaluate SegDAC on 8 ManiSkill3 manipulation tasks under 12 visual perturbation types across 3 difficulty levels. SegDAC improves over prior visual generalization methods by 15% on easy, 66% on medium, and 88% on the hardest settings. SegDAC matches the sample efficiency of the state-of-the-art visual RL methods while achieving improved generalization under visual changes.


Citation Information:

Alexandre Brown and Glen Berseth. "SegDAC: Visual Generalization in Reinforcement Learning via Dynamic Object Tokens." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

BibTeX:
@article{brown2026segdac,
    title={SegDAC: Visual Generalization in Reinforcement Learning via Dynamic Object Tokens},
    author={Alexandre Brown and Glen Berseth},
    journal={Reinforcement Learning Journal},
    volume={7},
    pages={},
    year={2026}
}