This is the pre-proceedings for the RLC 2026. You may expect minor changes.
Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.
Keywords: Policy Optimization, Average Reward.
We study model-free policy optimization algorithms for regret minimization in infinite horizon average reward Markov decision processes. Optimal regret bounds in this setting are only available under the restrictive uniform mixing assumption which states that, regardless of the policy one uses, the underlying Markov chain converges to its stationary behavior at a constant exponential rate. In this paper, we challenge this strong assumption and present the first provably efficient policy optimization method for infinite horizon average reward Markov decision processes without any mixing time dependence. Our algorithm achieves the optimal $\tilde{O}(\sqrt{T})$ regret assuming only a data coverage condition and the existence of a recurrent state that is re-visited within a finite expected time under all policies. Our results are obtained through a novel randomized episode length which allows us to leverage the Markov chain's renewal property.
William Powell, Jeongyeol Kwon, Qiaomin Xie, and Hanbaek Lyu. "Optimal Regret for Policy Optimization in Average Reward MDPs Without Mixing." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
BibTeX:@article{powell2026optimal,
title={Optimal Regret for Policy Optimization in Average Reward MDPs Without Mixing},
author={William Powell and Jeongyeol Kwon and Qiaomin Xie and Hanbaek Lyu},
journal={Reinforcement Learning Journal},
volume={7},
pages={},
year={2026}
}