This is the pre-proceedings for the RLC 2026. You may expect minor changes.
Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.
Keywords: Diversity, Robotics, Exploration
Being able to solve a task in diverse ways makes agents more robust to task variations and less prone to local optima. In this context, constrained diversity optimization has become a useful reinforcement learning (RL) framework for training a set of diverse agents in parallel. However, existing constrained-diversity RL methods often under-explore in complex tasks such as robot manipulation, resulting in limited behavioral diversity. We address this with a two-stage curriculum that introduces a spline-based trajectory prior as an inductive bias to produce diverse, high-reward behaviors in an initial stage, and then distills these behaviors into reactive, step-wise policies in a second stage. In our empirical evaluation, we provide novel insights into challenges of diversity-targeted training and show that our curriculum increases the diversity of learned skills while maintaining high task performance.
Cornelius V. Braun, Sayantan Auddy, and Marc Toussaint. "Trajectory First: A Curriculum for Discovering Diverse Policies." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
BibTeX:@article{braun2026trajectory,
title={Trajectory First: A Curriculum for Discovering Diverse Policies},
author={Cornelius V. Braun and Sayantan Auddy and Marc Toussaint},
journal={Reinforcement Learning Journal},
volume={7},
pages={},
year={2026}
}