This is the pre-proceedings for the RLC 2026. You may expect minor changes.
Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.
Keywords: Simulation-based planning, Stochastic tasks, Rollouts, Common random numbers,
Simulation-based planning with rollouts is a widely-deployed technique for decision making in stochastic environments. The primary instrument of simulation-based planning is a sampling model, which is repeatedly called to generate trajectories and estimate the utilities of available actions. Among the actions thus explored, one with the maximum estimated utility is then executed. In this paper, we examine the effect of using common random numbers in the simulation process. We obtain a simple recipe for (provably) reducing variance in relative utility when simulations invoke a rollout policy beyond some depth. Experiments on synthetic tasks confirm that our scheme improves task performance. The broader significance of our innovation is apparent from two practical applications: (1) single-step lookahead planning in a pension-disbursement task, and (2) a deployment of the well-known UCT algorithm for the game of Ludo.
Sandarbh Yadav, Frederic J Maliakkal, Harshad Khadilkar, and Shivaram Kalyanakrishnan. "Using Common Random Numbers for Simulation-based Planning with Rollouts." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
BibTeX:@article{yadav2026using,
title={Using Common Random Numbers for Simulation-based Planning with Rollouts},
author={Sandarbh Yadav and Frederic J Maliakkal and Harshad Khadilkar and Shivaram Kalyanakrishnan},
journal={Reinforcement Learning Journal},
volume={7},
pages={},
year={2026}
}