This is the pre-proceedings for the RLC 2026. You may expect minor changes.
Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.
Keywords: Foundations, philosophy, reward-is-not-enough, coherence
The Markov assumption is woven through the entire foundation of reinforcement learning (RL), allowing it to be cast as a purely instrumentalist framework of maximizing reward. However, deep RL agents *construct* their states, not necessarily ever reaching Markovness. Shifting out of theoretical inertia, we invite attention to the active process of state construction, away from the third-person bird's eye view into a first-person *phenomenological* perspective. We use this lens to challenge the primacy of reward maximization in non-Markov settings, arguing that it is not suited to account for the "spatial" or relational aspects of behaviour that become necessary. We take the Extended Mind thesis as a running thread, trace its evolution from functional equivalence to enactive sense making, and use it to inform our proposal. We propose that in non-Markov worlds another principle is necessary -- the drive for *coherence* -- a complementary coupling between the agent and the environment. To operationalize this beyond another auxiliary reward, we propose a radical but simple shift -- of the domain of the reward function from real to complex numbers. The imaginary dimension becomes a remarkably well-suited vehicle for expressing the philosophical concepts we explore, and naturally offers mathematics to express our coherence principle for the now-complex representation. This work sketches the foundations for a new generation of RL agents that are concerned with more than their gain, and offers a completely new lens for considering classical open problems.
Anna Harutyunyan, Will Dabney, and Doina Precup. "Cohering Reinforcement Learning." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
BibTeX:@article{harutyunyan2026cohering,
title={Cohering Reinforcement Learning},
author={Anna Harutyunyan and Will Dabney and Doina Precup},
journal={Reinforcement Learning Journal},
volume={7},
pages={},
year={2026}
}