This is the pre-proceedings for the RLC 2026. You may expect minor changes.
Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.
Keywords: Bayesian Q-learning, distributional reinforcement learning, risk, exploration, value of
Distributional reinforcement learning (DRL) conventionally models objective uncertainty in the return distributions from states, allowing important quantities to be computed such as risk-sensitive policies. Bayesian approaches additionally quantify subjective, or epistemic, uncertainty and can therefore provide principled signals for exploration. However, most Bayesian value-based reinforcement learning algorithms, such as the seminal Bayesian Q-learning (BQL), rely on restrictive likelihood assumptions to remain tractable. This severely constrains the set of approximating return distributions, and thus the risk-relevant statistics they can represent. Due to these limitations, Bayesian risk-sensitive exploration has been relatively underexplored. We propose a Bayesian DRL algorithm which maintains and updates a full posterior belief over finite-dimensional mean embeddings of the return distribution, chosen to capture higher-order moments and admit Bellman-consistent propagation. This representation allows us to generalize the value of perfect information exploration policy that BQL employs to moment-based risk objectives, enabling exploration that explicitly trades off epistemic uncertainty and risk. As we show in simulations, our algorithm learns more accurate approximations to return distributions compared to BQL, and the resulting risk-sensitive exploration improves control performance over a matched non-Bayesian DRL baseline.
Karim Zaghw, Peter Dayan, and Georgy Antonov. "Weight a moment: risk-aware exploration with Bayes." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
BibTeX:@article{zaghw2026weight,
title={Weight a moment: risk-aware exploration with Bayes},
author={Karim Zaghw and Peter Dayan and Georgy Antonov},
journal={Reinforcement Learning Journal},
volume={7},
pages={},
year={2026}
}