This is the pre-proceedings for the RLC 2026. You may expect minor changes.

Hierarchical Behaviour Spaces

By Michael Matthews, Pierluca D'Oro, Anssi Kanervisto, Scott Fujimoto, Jakob Nicolaus Foerster, and Mikael Henaff

Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.


Download:

Keywords: Hierarchical RL, Exploration

Abstract:

Recent work in hierarchical reinforcement learning has shown success in scaling to billions of timesteps when learning over a set of predefined option reward functions. We show that, instead of using a single reward function per option, the reward functions can be effectively used to induce a space of behaviours, by letting the controller specify linear combinations over reward functions, allowing a more expressive set of policies to be represented. We call this method Hierarchical Behaviour Spaces (HBS). We evaluate HBS on the NetHack Learning Environment, demonstrating strong performance. We conduct a series of experiments and determine that, perhaps going against conventional wisdom, the benefits of hierarchy in our method come from increased exploration rather than long term reasoning.


Citation Information:

Michael Matthews, Pierluca D'Oro, Anssi Kanervisto, Scott Fujimoto, Jakob Nicolaus Foerster, and Mikael Henaff. "Hierarchical Behaviour Spaces." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

BibTeX:
@article{matthews2026hierarchical,
    title={Hierarchical Behaviour Spaces},
    author={Michael Matthews and Pierluca D'Oro and Anssi Kanervisto and Scott Fujimoto and Jakob Nicolaus Foerster and Mikael Henaff},
    journal={Reinforcement Learning Journal},
    volume={7},
    pages={},
    year={2026}
}