This is the pre-proceedings for the RLC 2026. You may expect minor changes.
Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
Will be presented at the Reinforcement Learning Conference (RLC), Montréal, Quebec, Canada, August 15–17, 2026.
Keywords: Deep Reinforcement Learning, On-Policy, Proximal Policy Optimization
On-policy Reinforcement Learning (RL) offers desirable features such as stable learning, fewer policy updates, and the ability to evaluate a policy’s return during training. While recent efforts have focused on off-policy methods, achieving significant advancements, PPO remains the go-to algorithm for on-policy RL due to its apparent simplicity and effectiveness. Nonetheless, PPO relies on subtle, often poorly documented adjustments that can critically affect its performance. Our proposed approach, PPO+, is a principled adaptation of the PPO algorithm that strengthens its adherence to the on-policy objective, enhancing stability and efficiency. PPO+ demonstrates significantly improved asymptotic performance over PPO, and a substantially reduced performance gap with off-policy algorithms in several challenging continuous control tasks. Beyond just performance, our findings offer a fresh perspective on on-policy RL that can serve as guidance for future research endeavors.
Mahdi Kallel, Jose-Luis Holgado-Alvarez, Samuele Tosatto, and Carlo D'Eramo. "PPO+: Enhancing proximal policy optimization." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
BibTeX:@article{kallel2026ppo,
title={PPO+: Enhancing proximal policy optimization},
author={Mahdi Kallel and Jose-Luis Holgado-Alvarez and Samuele Tosatto and Carlo D'Eramo},
journal={Reinforcement Learning Journal},
volume={7},
pages={},
year={2026}
}