This is the pre-proceedings for the RLC 2026. You may expect minor changes.

Direct Advantage Estimation for Scalable and Sample-efficient Deep Reinforcement Learning

By Hsiao-Ru Pan, and Bernhard Schölkopf

Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

Will be presented at the Reinforcement Learning Conference (RLC), Montréal, Quebec, Canada, August 15–17, 2026.


Download:

Keywords: advantage estimation, sample efficient deep RL, POMDP, multi-step learning

Abstract:

Direct Advantage Estimation (DAE) has been shown to improve the sample efficiency of deep reinforcement learning algorithms. However, its reliance on full environment observability limits its applicability in realistic settings, and its requirement to model transition probabilities incurs substantial computational overhead for high-dimensional observations. In the present work, we address both limitations. First, we extend the theoretical framework of DAE to partially observable domains with minimal modifications. Second, we reduce its computational complexity by introducing discrete latent dynamics models that efficiently approximate transition probabilities. We evaluate our approach on the Arcade Learning Environment and find that DAE scales effectively with function approximator capacity while retaining high sample efficiency.


Citation Information:

Hsiao-Ru Pan and Bernhard Schölkopf. "Direct Advantage Estimation for Scalable and Sample-efficient Deep Reinforcement Learning." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

BibTeX:
@article{pan2026direct,
    title={Direct Advantage Estimation for Scalable and Sample-efficient Deep Reinforcement Learning},
    author={Hsiao-Ru Pan and Bernhard Schölkopf},
    journal={Reinforcement Learning Journal},
    volume={7},
    pages={},
    year={2026}
}