This is the pre-proceedings for the RLC 2026. You may expect minor changes.

Cohering Reinforcement Learning

By Anna Harutyunyan, Will Dabney, and Doina Precup

Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.


Download:

Keywords: Foundations, philosophy, reward-is-not-enough, coherence

Abstract:

The Markov assumption is woven through the entire foundation of reinforcement learning (RL), allowing it to be cast as a purely instrumentalist framework of maximizing reward. However, deep RL agents *construct* their states, not necessarily ever reaching Markovness. Shifting out of theoretical inertia, we invite attention to the active process of state construction, away from the third-person bird's eye view into a first-person *phenomenological* perspective. We use this lens to challenge the primacy of reward maximization in non-Markov settings, arguing that it is not suited to account for the "spatial" or relational aspects of behaviour that become necessary. We take the Extended Mind thesis as a running thread, trace its evolution from functional equivalence to enactive sense making, and use it to inform our proposal. We propose that in non-Markov worlds another principle is necessary -- the drive for *coherence* -- a complementary coupling between the agent and the environment. To operationalize this beyond another auxiliary reward, we propose a radical but simple shift -- of the domain of the reward function from real to complex numbers. The imaginary dimension becomes a remarkably well-suited vehicle for expressing the philosophical concepts we explore, and naturally offers mathematics to express our coherence principle for the now-complex representation. This work sketches the foundations for a new generation of RL agents that are concerned with more than their gain, and offers a completely new lens for considering classical open problems.


Citation Information:

Anna Harutyunyan, Will Dabney, and Doina Precup. "Cohering Reinforcement Learning." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

BibTeX:
@article{harutyunyan2026cohering,
    title={Cohering Reinforcement Learning},
    author={Anna Harutyunyan and Will Dabney and Doina Precup},
    journal={Reinforcement Learning Journal},
    volume={7},
    pages={},
    year={2026}
}