This is the pre-proceedings for the RLC 2026. You may expect minor changes.

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification

By Haoyang Hong, Zichen Wang, Quanquan Gu, and Huazheng Wang

Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.


Download:

Keywords: KL regularization, contextual bandits, reinforcement learning, model misspecification

Abstract:

We study KL-regularized contextual bandits and episodic reinforcement learning (RL) under general function approximation with model misspecification. Existing guarantees rely on realizability and therefore do not extend to misspecified models, where classical regret bounds may fail. This work introduces KL misspecification formulations for contextual bandits and episodic RL and analyzes regression-based algorithms that via Gibbs policies updates. High-probability KL-regret guarantees with explicit misspecification terms are established, recovering the standard realizable KL-regularized setting as a special case.


Citation Information:

Haoyang Hong, Zichen Wang, Quanquan Gu, and Huazheng Wang. "Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

BibTeX:
@article{hong2026online,
    title={Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification},
    author={Haoyang Hong and Zichen Wang and Quanquan Gu and Huazheng Wang},
    journal={Reinforcement Learning Journal},
    volume={7},
    pages={},
    year={2026}
}