This is the pre-proceedings for the RLC 2026. You may expect minor changes.
Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.
Keywords: multi-policy, multi objective reinforcement learning, infinite-horizon.
Multi-objective reinforcement learning (MORL) addresses problems with multiple, often conflicting goals by seeking a set of trade-off policies rather than a single solution. Existing tabular solutions are limited to episodic problems and are incompatible with infinite-horizon settings. In this work, we address this gap by deriving a set of design principles for tabular multi-policy MORL, where the proposed framework supports both stationary and non-stationary policies, prevents spurious domination, and incorporates cycle detection for robust return guarantees. Through ablation studies, we show how each principle contributes to discovering diverse and reliable policies that map the entire Pareto front.
Marcelo d'Almeida and Daniel Mosse. "Design Principles for Tabular Multi-Policy MORL in Infinite Horizons." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
BibTeX:@article{dalmeida2026design,
title={Design Principles for Tabular Multi-Policy MORL in Infinite Horizons},
author={Marcelo d'Almeida and Daniel Mosse},
journal={Reinforcement Learning Journal},
volume={7},
pages={},
year={2026}
}