This is the pre-proceedings for the RLC 2026. You may expect minor changes.

PPO+: Enhancing proximal policy optimization

By Mahdi Kallel, Jose-Luis Holgado-Alvarez, Samuele Tosatto, and Carlo D'Eramo

Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

Will be presented at the Reinforcement Learning Conference (RLC), Montréal, Quebec, Canada, August 15–17, 2026.


Download:

Keywords: Deep Reinforcement Learning, On-Policy, Proximal Policy Optimization

Abstract:

On-policy Reinforcement Learning (RL) offers desirable features such as stable learning, fewer policy updates, and the ability to evaluate a policy’s return during training. While recent efforts have focused on off-policy methods, achieving significant advancements, PPO remains the go-to algorithm for on-policy RL due to its apparent simplicity and effectiveness. Nonetheless, PPO relies on subtle, often poorly documented adjustments that can critically affect its performance. Our proposed approach, PPO+, is a principled adaptation of the PPO algorithm that strengthens its adherence to the on-policy objective, enhancing stability and efficiency. PPO+ demonstrates significantly improved asymptotic performance over PPO, and a substantially reduced performance gap with off-policy algorithms in several challenging continuous control tasks. Beyond just performance, our findings offer a fresh perspective on on-policy RL that can serve as guidance for future research endeavors.


Citation Information:

Mahdi Kallel, Jose-Luis Holgado-Alvarez, Samuele Tosatto, and Carlo D'Eramo. "PPO+: Enhancing proximal policy optimization." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

BibTeX:
@article{kallel2026ppo,
    title={PPO+: Enhancing proximal policy optimization},
    author={Mahdi Kallel and Jose-Luis Holgado-Alvarez and Samuele Tosatto and Carlo D'Eramo},
    journal={Reinforcement Learning Journal},
    volume={7},
    pages={},
    year={2026}
}