This is the pre-proceedings for the RLC 2026. You may expect minor changes.

When Can Pure Exploitation Succeed in Linear RL? Decoys and Self-Identifiability for Greedy LSVI

By Manoj Saravanan, and Rohit Kumar Salla

Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.


Download:

Keywords: theoretical reinforcement learning; exploration-free reinforcement learning; linear

Abstract:

We study online episodic reinforcement learning in finite-horizon Markov decision processes with a finite action set, known features, bounded rewards, and linear Bellman completeness. We focus on Greedy-LSVI, which performs stage-wise ridge regression on Bellman targets followed by deterministic greedy action selection, with no optimism or randomization. We identify a precise information-theoretic obstruction to exploration-free learning: a decoy, namely an alternative environment in the same model class that induces the same trajectory law under a greedy-representable policy $\pi$ while making $\pi$ globally optimal, even though $\pi$ is suboptimal in the true environment. We call an instance self-identifiable if no suboptimal greedy-representable policy admits a decoy. Our main result shows that, up to a mild model-class closure assumption, decoys are equivalent to the existence of an on-policy covariance null direction that is still action-sensitive on states visited by $\pi$. This yields a dichotomy for Greedy-LSVI: decoys permit linear-regret instances, whereas under self-identifiability and additional bounded-parameter and leverage conditions, Greedy-LSVI achieves $\widetilde{O}(\sqrt{K})$ regret over $K$ episodes and, under a witnessed-margin condition, $\widetilde{O}(\log K)$ regret. We also construct a self-identifiable instance with rank-deficient on-policy covariance, showing that full-rank covariance is not necessary for pure exploitation to succeed.


Citation Information:

Manoj Saravanan and Rohit Kumar Salla. "When Can Pure Exploitation Succeed in Linear RL? Decoys and Self-Identifiability for Greedy LSVI." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

BibTeX:
@article{saravanan2026when,
    title={When Can Pure Exploitation Succeed in Linear RL? Decoys and Self-Identifiability for Greedy LSVI},
    author={Manoj Saravanan and Rohit Kumar Salla},
    journal={Reinforcement Learning Journal},
    volume={7},
    pages={},
    year={2026}
}