This is the pre-proceedings for the RLC 2026. You may expect minor changes.

Trajectory First: A Curriculum for Discovering Diverse Policies

By Cornelius V. Braun, Sayantan Auddy, and Marc Toussaint

Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.


Download:

Keywords: Diversity, Robotics, Exploration

Abstract:

Being able to solve a task in diverse ways makes agents more robust to task variations and less prone to local optima. In this context, constrained diversity optimization has become a useful reinforcement learning (RL) framework for training a set of diverse agents in parallel. However, existing constrained-diversity RL methods often under-explore in complex tasks such as robot manipulation, resulting in limited behavioral diversity. We address this with a two-stage curriculum that introduces a spline-based trajectory prior as an inductive bias to produce diverse, high-reward behaviors in an initial stage, and then distills these behaviors into reactive, step-wise policies in a second stage. In our empirical evaluation, we provide novel insights into challenges of diversity-targeted training and show that our curriculum increases the diversity of learned skills while maintaining high task performance.


Citation Information:

Cornelius V. Braun, Sayantan Auddy, and Marc Toussaint. "Trajectory First: A Curriculum for Discovering Diverse Policies." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

BibTeX:
@article{braun2026trajectory,
    title={Trajectory First: A Curriculum for Discovering Diverse Policies},
    author={Cornelius V. Braun and Sayantan Auddy and Marc Toussaint},
    journal={Reinforcement Learning Journal},
    volume={7},
    pages={},
    year={2026}
}