This is the pre-proceedings for the RLC 2026. You may expect minor changes.
Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.
Keywords: Reinforcement learning, decision trees, MDPs, POMDPs
For applications like medicine, machine learning models ought to be interpretable. In that case, decision tree models are preferred over neural networks because humans can read their predictions from the root to the leaves. Training such decision trees for sequential decision making problems is a relatively new research direction and most of the existing literature focuses on imitating neural networks. In contrast, we study reinforcement learning (RL) algorithms that train decision trees directly optimizing some trade-off of cumulative rewards and interpretability in a Markov decision process (MDP). We show that such algorithms can be seen as training policies for partially observable Markov decision processes (POMDPs). This helps us understand why in practice it is often hard to train decision tree policies from scratch in MDPs.
Hector Kohler, Riad Akrour, and Philippe Preux. "Limits of reinforcement learning for decision trees in Markov decision processes." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
BibTeX:@article{kohler2026limits,
title={Limits of reinforcement learning for decision trees in Markov decision processes},
author={Hector Kohler and Riad Akrour and Philippe Preux},
journal={Reinforcement Learning Journal},
volume={7},
pages={},
year={2026}
}