This is the pre-proceedings for the RLC 2026. You may expect minor changes.
Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
Will be presented at the Reinforcement Learning Conference (RLC), Montréal, Quebec, Canada, August 15–17, 2026.
Keywords: Diffusion policies, Efficient Computing, Prefix Value Functions, Continuous Control
Diffusion policies generate actions through a K-step denoising chain, but standard training optimizes only the terminal output, leaving intermediate iterates unsupervised and nonexecutable. We introduce POGP (Prefix-Optimal Generative Policies), which learns a prefix value function Vt(s,at) at every denoising step via a Bellman-style recursion over the chain. The PVF serves a dual purpose: as an auxiliary training objective it improves intermediate action quality, and as a test-time metric the marginal change across consecutive steps provides a parameter-free stopping criterion. On four MuJoCo environments against 12 baselines, POGP outperforms state-of-the-art methods dynamic diffusion baselines by approximately 3.5% while reducing denoising iterations by ≈2.7×compared to fixed-step diffusion baselines, while retaining similar perfomrance gains.
Rohit Kumar Salla, Manoj Saravanan, and Simon Stepputtis. "Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
BibTeX:@article{salla2026learning,
title={Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control},
author={Rohit Kumar Salla and Manoj Saravanan and Simon Stepputtis},
journal={Reinforcement Learning Journal},
volume={7},
pages={},
year={2026}
}