This is the pre-proceedings for the RLC 2026. You may expect minor changes.

On the Variance of Temporal Difference Learning and its Reduction Using Control Variates

By Hsiao-Ru Pan, and Bernhard Schölkopf

Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

Will be presented at the Reinforcement Learning Conference (RLC), Montréal, Quebec, Canada, August 15–17, 2026.


Download:

Keywords: temporal difference learning, Monte Carlo estimation, advantage function, control

Abstract:

We analyze the variance of temporal difference (TD) learning using the phased setting with tabular representation, and show that one of the mechanisms behind its ability to reduce variance is by effectively aggregating over a larger number of independent trajectories. Based on this insight, we demonstrate that (1) the variance of TD is asymptotically bounded from above by Monte Carlo (MC) estimators, and (2) shorter horizon updates incurs less variance for a fixed number of samples. Beyond TD, we show that Direct Advantage Estimation (DAE), a method for estimating the advantage function, can be seen as a type of regression-adjusted control variate, which achieves a tighter bound on the variance compared to TD in the large-sample limit. Finally, we numerically illustrate the behaviors of these estimators with carefully designed environments.


Citation Information:

Hsiao-Ru Pan and Bernhard Schölkopf. "On the Variance of Temporal Difference Learning and its Reduction Using Control Variates." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

BibTeX:
@article{pan2026on,
    title={On the Variance of Temporal Difference Learning and its Reduction Using Control Variates},
    author={Hsiao-Ru Pan and Bernhard Schölkopf},
    journal={Reinforcement Learning Journal},
    volume={7},
    pages={},
    year={2026}
}