This is the pre-proceedings for the RLC 2026. You may expect minor changes.
Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
Will be presented at the Reinforcement Learning Conference (RLC), Montréal, Quebec, Canada, August 15–17, 2026.
Keywords: Risk-averse Reinforcement Learning, Sample Complexity, Optimized Certainty
We study risk-sensitive reinforcement learning in tabular discounted Markov Decision Processes (MDPs), assuming access to a generative model of the MDP. We consider a family of risk measures called the optimized certainty equivalents (OCEs), which includes important risk measures such as entropic risk, CVaR, and the mean-variance criterion. Our focus is on the sample complexities of learning the optimal state–action value function (value learning) and an optimal policy (policy learning) under recursive OCEs. We provide an exact characterization of the utility functions $u$ for which the corresponding recursive OCE objective is PAC-learnable. We establish that whenever $u$ does not have full domain, i.e., $\operatorname{dom}(u)\neq \mathbb{R}$, the corresponding problem is not PAC-learnable. We present and analyze a simple model-based algorithm and derive PAC sample complexity bounds for both value and policy learning. Finally, we establish complementary lower bounds, demonstrating tightness with respect to the size $SA$ of the state-action space. For a more restricted class of utilities, we derive lower bounds that make the dependence on the effective horizon, $1/(1-\gamma)$, explicit. In particular, for $\text{CVaR}_\tau$, our lower bound scales as $\tau^{-2}$, improving the dependence on $\tau$ by a factor of $\tau^{-1}$ over the best existing lower bound, although its dependence on $1/(1-\gamma)$ remains suboptimal.
Oliver Mortensen and M. Sadegh Talebi. "On the Sample Complexity of Discounted Reinforcement Learning with Optimized Certainty Equivalents." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
BibTeX:@article{mortensen2026on,
title={On the Sample Complexity of Discounted Reinforcement Learning with Optimized Certainty Equivalents},
author={Oliver Mortensen and M. Sadegh Talebi},
journal={Reinforcement Learning Journal},
volume={7},
pages={},
year={2026}
}