This is the pre-proceedings for the RLC 2026. You may expect minor changes.
Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
Will be presented at the Reinforcement Learning Conference (RLC), Montréal, Quebec, Canada, August 15–17, 2026.
Keywords: Reinforcement Learning, Partial Observability, Distributional Reinforcement Learning,
In many real-world planning tasks, agents must tackle uncertainty about the environment’s state and variability in the outcomes induced by stochastic dynamics and rewards. Motivated by recent progress in world model approaches---where latent models approximate beliefs and support planning---we extend Distributional Reinforcement Learning (DistRL), which models the entire return distribution for fully observable domains, to Partially Observable Markov Decision Processes (POMDPs). Concretely, we introduce new distributional Bellman operators for partial observability, prove contraction of the evaluation operator, and establish convergence of the optimality operator iterates under a unique optimal policy assumption in the supremum $p$-Wasserstein metric. We also propose a finite representation of these return distributions via $\psi$-vectors, generalizing the classical $\alpha$-vectors in POMDP solvers. Building on this, we develop Distributional Point-Based Value Iteration (DPBVI), which integrates $\psi$-vectors into a standard point-based backup procedure—bridging DistRL and POMDP planning. Our experiments demonstrate that DPBVI recovers classical Point-Based Value Iteration (PBVI) in the risk-neutral case, validating the distributional extension.
Larry Preuett, Qiuyi Zhang, and Muhammad Aurangzeb Ahmad. "Provable Distributional Value Iteration under Partial Observability." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
BibTeX:@article{preuett2026provable,
title={Provable Distributional Value Iteration under Partial Observability},
author={Larry Preuett and Qiuyi Zhang and Muhammad Aurangzeb Ahmad},
journal={Reinforcement Learning Journal},
volume={7},
pages={},
year={2026}
}