This is the pre-proceedings for the RLC 2026. You may expect minor changes.

Fully Offline Reinforcement Learning

By Mattie Fellows, Clarisse Wibault, Uljad Berdica, Johannes Forkel, Michael A Osborne, and Jakob Nicolaus Foerster

Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.


Download:

Keywords: Offline RL, Bayesian RL, Model-Based ORL, Regret Analysis

Abstract:

Sample efficiency remains a major barrier to the real-world deployment of Reinforcement Learning (RL), where large numbers of online interactions are often costly or unsafe. Offline RL (ORL) seeks to address this challenge by learning policies from static datasets, yet existing methods rely on undocumented online interactions for hyperparameter tuning. Moreover, current methods provide no reliable estimate of their initial online performance using offline data alone. To overcome these limitations, we introduce two complementary algorithms for fully offline RL. Firstly, SOReL (Safe Offline Reinforcement Learning) is a model-based Bayesian approach that infers a posterior over environment dynamics from offline data and learns an approximation to a Bayes-optimal policy entirely offline. By quantifying predictive uncertainty via model ensembling, SOReL enables reliable and tractable offline prediction of online policy value and fully offline hyperparameter selection. Secondly, TOReL (Tuning for Offline Reinforcement Learning) extends this principle to arbitrary model-free and non-Bayesian ORL algorithms, leveraging predictive uncertainty quantification to eliminate costly online tuning loops. We additionally provide a regret analysis of offline Bayesian RL, showing that Bayesian approaches achieve the optimal frequentist minimax regret rate, thereby offering a formal frequentist justification for offline Bayesian methods. Empirical results demonstrate that SOReL accurately predicts online performance and that, using only offline data, TOReL achieves competitive results with online hyperparameter tuning methods, advancing safe and reliable fully offline RL for real-world deployment. Our implementations are publicly available.


Citation Information:

Mattie Fellows, Clarisse Wibault, Uljad Berdica, Johannes Forkel, Michael A Osborne, and Jakob Nicolaus Foerster. "Fully Offline Reinforcement Learning." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

BibTeX:
@article{fellows2026fully,
    title={Fully Offline Reinforcement Learning},
    author={Mattie Fellows and Clarisse Wibault and Uljad Berdica and Johannes Forkel and Michael A Osborne and Jakob Nicolaus Foerster},
    journal={Reinforcement Learning Journal},
    volume={7},
    pages={},
    year={2026}
}