This is the pre-proceedings for the RLC 2026. You may expect minor changes.
Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.
Keywords: KL regularization, contextual bandits, reinforcement learning, model misspecification
We study KL-regularized contextual bandits and episodic reinforcement learning (RL) under general function approximation with model misspecification. Existing guarantees rely on realizability and therefore do not extend to misspecified models, where classical regret bounds may fail. This work introduces KL misspecification formulations for contextual bandits and episodic RL and analyzes regression-based algorithms that via Gibbs policies updates. High-probability KL-regret guarantees with explicit misspecification terms are established, recovering the standard realizable KL-regularized setting as a special case.
Haoyang Hong, Zichen Wang, Quanquan Gu, and Huazheng Wang. "Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
BibTeX:@article{hong2026online,
title={Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification},
author={Haoyang Hong and Zichen Wang and Quanquan Gu and Huazheng Wang},
journal={Reinforcement Learning Journal},
volume={7},
pages={},
year={2026}
}