This is the pre-proceedings for the RLC 2026. You may expect minor changes.

Risk-Aware General-Utility Markov Decision Processes

By Pedro Pinto Santos, Fábio Vital, Alberto Sardinha, and Francisco S. Melo

Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

Will be presented at the Reinforcement Learning Conference (RLC), Montréal, Quebec, Canada, August 15–17, 2026.


Download:

Keywords: Risk-aware decision-making, Reinforcement learning, Planning.

Abstract:

We study general-utility Markov decision processes (GUMDPs) with risk-aware objectives. In this framework, an agent aims to optimize a risk measure of the distribution of objective values, where the objective function depends on the frequency of visitation of states induced by the agent's policy. First, we motivate, propose, and formalize risk-aware GUMDPs, which enable agents and decision makers to trade off expected performance by risk aversion while benefiting from the rich set of objectives that can be cast under the framework of GUMDPs. We focus our attention on the entropic risk measure (ERM). Second, we show how we can solve risk-aware GUMDPs with ERM objectives by resorting to online planning techniques. In particular, we propose an approach based on Monte Carlo Tree Search (MCTS) to provably solve risk-aware GUMDPs up to any desired accuracy. Third, we provide a set of experimental results showcasing that our approach is successful when optimizing for a spectrum of risk-aware behaviors in the context of GUMDPs under diverse tasks (standard MDPs, maximum state entropy exploration, imitation learning, and multi-objective MDPs).


Citation Information:

Pedro Pinto Santos, Fábio Vital, Alberto Sardinha, and Francisco S. Melo. "Risk-Aware General-Utility Markov Decision Processes." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

BibTeX:
@article{santos2026riskaware,
    title={Risk-Aware General-Utility Markov Decision Processes},
    author={Pedro Pinto Santos and Fábio Vital and Alberto Sardinha and Francisco S. Melo},
    journal={Reinforcement Learning Journal},
    volume={7},
    pages={},
    year={2026}
}