This is the pre-proceedings for the RLC 2026. You may expect minor changes.
Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
Will be presented at the Reinforcement Learning Conference (RLC), Montréal, Quebec, Canada, August 15–17, 2026.
Keywords: Curriculum Reinforcement Learning, Distributionally Robust Reinforcement Learning,
A central challenge in reinforcement learning is that policies trained in controlled environments often fail under distribution shifts at deployment into real-world environments. Distributionally Robust Reinforcement Learning (DRRL) addresses this by optimizing for worst-case performance within an uncertainty set defined by a robustness budget $\epsilon$. However, fixing $\epsilon$ results in a tradeoff between performance and robustness: small values yield high nominal performance but weak robustness, while large values can result in instability and overly conservative policies. We propose Distributionally Robust Self-Paced Curriculum Reinforcement Learning (DR-SPCRL), a method that overcomes this limitation by treating $\epsilon$ as a continuous curriculum. DR-SPCRL adaptively schedules the robustness budget according to the agent’s progress, enabling a balance between nominal and robust performance. Empirical results across multiple environments demonstrate that DR-SPCRL not only stabilizes training but also achieves a superior robustness–performance trade-off, yielding an average 24.1\% increase in episodic return under varying perturbations compared to fixed or heuristic scheduling strategies.
Anirudh Satheesh, Keenan Powell, and Vaneet Aggarwal. "Distributionally Robust Self Paced Curriculum Reinforcement Learning." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
BibTeX:@article{satheesh2026distributionally,
title={Distributionally Robust Self Paced Curriculum Reinforcement Learning},
author={Anirudh Satheesh and Keenan Powell and Vaneet Aggarwal},
journal={Reinforcement Learning Journal},
volume={7},
pages={},
year={2026}
}