This is the pre-proceedings for the RLC 2026. You may expect minor changes.

Goal-Oriented Reinforcement Learning for Stochastic Shortest Paths with Dead-Ends

By Gustavo De Mari Pereira, and Leliane N. de Barros

Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.


Download:

Keywords: Reinforcement Learning, Stochastic Shortest Path, Dead-ends, Model-based RL, Dual

Abstract:

Stochastic Shortest Path (SSP) problems provide a natural framework for goal-oriented reinforcement learning. However, the classical SSP assumption that a proper policy exists (i.e. a policy that reaches the goal with probability 1) can be violated in practice by the presence of dead-ends. When the goal is unreachable from some states, standard undiscounted RL may diverge, while discounted RL can lead to suboptimal goal-reaching behavior. This paper addresses these challenges by introducing Finite-Penalty Q-learning and Dual Dyna-Q, two algorithms designed to handle SSPs with dead-ends. We evaluate our methods across a variety of grid navigation and symbolic planning (Probabilistic PDDL) benchmarks, demonstrating that they reliably reach the goal while maintaining cost efficiency, significantly outperforming standard model-free and discounted baseline solutions in environments with dead-ends.


Citation Information:

Gustavo De Mari Pereira and Leliane N. de Barros. "Goal-Oriented Reinforcement Learning for Stochastic Shortest Paths with Dead-Ends." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

BibTeX:
@article{pereira2026goaloriented,
    title={Goal-Oriented Reinforcement Learning for Stochastic Shortest Paths with Dead-Ends},
    author={Gustavo De Mari Pereira and Leliane N. de Barros},
    journal={Reinforcement Learning Journal},
    volume={7},
    pages={},
    year={2026}
}