This is the pre-proceedings for the RLC 2026. You may expect minor changes.
Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.
Keywords: Reinforcement Learning, Stochastic Shortest Path, Dead-ends, Model-based RL, Dual
Stochastic Shortest Path (SSP) problems provide a natural framework for goal-oriented reinforcement learning. However, the classical SSP assumption that a proper policy exists (i.e. a policy that reaches the goal with probability 1) can be violated in practice by the presence of dead-ends. When the goal is unreachable from some states, standard undiscounted RL may diverge, while discounted RL can lead to suboptimal goal-reaching behavior. This paper addresses these challenges by introducing Finite-Penalty Q-learning and Dual Dyna-Q, two algorithms designed to handle SSPs with dead-ends. We evaluate our methods across a variety of grid navigation and symbolic planning (Probabilistic PDDL) benchmarks, demonstrating that they reliably reach the goal while maintaining cost efficiency, significantly outperforming standard model-free and discounted baseline solutions in environments with dead-ends.
Gustavo De Mari Pereira and Leliane N. de Barros. "Goal-Oriented Reinforcement Learning for Stochastic Shortest Paths with Dead-Ends." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
BibTeX:@article{pereira2026goaloriented,
title={Goal-Oriented Reinforcement Learning for Stochastic Shortest Paths with Dead-Ends},
author={Gustavo De Mari Pereira and Leliane N. de Barros},
journal={Reinforcement Learning Journal},
volume={7},
pages={},
year={2026}
}