This is the pre-proceedings for the RLC 2026. You may expect minor changes.

Statistical Inference for Policy Evaluation with Temporal Difference Learning

By Weichen Wu, Gen Li, Yuting Wei, and Alessandro Rinaldo

Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.


Download:

Keywords: Temporal Difference Learning, Statistical Inference, Berry-Esseen bound, Central Limit

Abstract:

We investigate the statistical properties of Temporal Difference (TD) learning with Polyak-Ruppert averaging, arguably one of the most widely used algorithms in reinforcement learning, for the task of estimating the parameters of the optimal linear approximation to the value function. Assuming independent samples, we make three theoretical contributions that improve upon the current state-of-the-art results: (i) we establish refined high-dimensional Berry-Esseen bounds over the class of convex sets, achieving faster rates than the best known results, and (ii) we propose and analyze a novel, computationally efficient online plug-in estimator of the asymptotic covariance matrix; (iii) we derive sharper high probability convergence guarantees that depend explicitly on the asymptotic variance and hold under weaker conditions than those adopted in the literature. These results enable the construction of confidence regions and simultaneous confidence intervals for the linear parameters of the value function approximation, with guaranteed finite-sample coverage. We demonstrate the applicability of our theoretical findings through numerical experiments.


Citation Information:

Weichen Wu, Gen Li, Yuting Wei, and Alessandro Rinaldo. "Statistical Inference for Policy Evaluation with Temporal Difference Learning." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

BibTeX:
@article{wu2026statistical,
    title={Statistical Inference for Policy Evaluation with Temporal Difference Learning},
    author={Weichen Wu and Gen Li and Yuting Wei and Alessandro Rinaldo},
    journal={Reinforcement Learning Journal},
    volume={7},
    pages={},
    year={2026}
}