This is the pre-proceedings for the RLC 2026. You may expect minor changes.

Q-Based Variational Inverse Reinforcement Learning

By Ondrej Bajgar, Peter Tisnikar, Konstantinos Gatsis, Alessandro Abate, and Michael A Osborne

Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.


Download:

Keywords: inverse reinforcement learning, Bayesian IRL, variational inference, uncertainty

Abstract:

The development of safe and beneficial AI requires that systems can learn and act in accordance with human preferences. However, explicitly specifying these preferences by hand is often infeasible. Inverse reinforcement learning (IRL) addresses this challenge by inferring preferences, represented as reward functions, from expert behaviour. We introduce Q-based Variational IRL (QVIRL), a novel Bayesian IRL method that recovers a posterior distribution over rewards from expert demonstrations via primarily learning a variational distribution over optimal Q-values. Unlike previous approaches, QVIRL combines scalability with uncertainty quantification, important for safety-critical applications as well as active learning. We demonstrate QVIRL's strong performance in apprenticeship learning across various tasks, including gridworlds, Lunar Lander, the Highway Environment, and two ATARI games both with static expert data and with active learning. It is the first method for Bayesian IRL that demonstrates training from raw pixel observations.


Citation Information:

Ondrej Bajgar, Peter Tisnikar, Konstantinos Gatsis, Alessandro Abate, and Michael A Osborne. "Q-Based Variational Inverse Reinforcement Learning." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

BibTeX:
@article{bajgar2026qbased,
    title={Q-Based Variational Inverse Reinforcement Learning},
    author={Ondrej Bajgar and Peter Tisnikar and Konstantinos Gatsis and Alessandro Abate and Michael A Osborne},
    journal={Reinforcement Learning Journal},
    volume={7},
    pages={},
    year={2026}
}