This is the pre-proceedings for the RLC 2026. You may expect minor changes.
Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.
Keywords: Human-AI Interaction, Chess, Sequential Decision Making, AI Interventions
Artificial intelligence systems are increasingly used to assist humans in sequential decision-making tasks. A common approach is to recommend actions that are optimal according to a strong model of the environment. Such recommendations, however, implicitly assume that the decision maker will continue to act optimally afterward. Human decision makers often deviate from optimal play, meaning that actions that are optimal in isolation may lead to states that are difficult for humans to navigate successfully. This raises a fundamental question: how should AI systems intervene when assisting suboptimal decision makers in sequential environments? In this work, we study value-aware interventions that account for how humans are likely to behave after receiving assistance. Our approach builds on a basic principle from reinforcement learning: when a policy is suboptimal, discrepancies arise between the actions taken by the policy and those that maximize the expected value of future outcomes. We show how such discrepancies can be used to identify opportunities where interventions improve downstream performance. To operationalize this idea, we develop a data-driven framework that learns models of human decision-making from large datasets and uses these models to guide intervention decisions. We study these ideas in the domain of chess, which provides large-scale human gameplay data and strong AI systems for evaluating intervention strategies.
Saumik Narayanan, Raja Panjwani, Siddhartha Sen, and Chien-Ju Ho. "Improving Human Performance with Value-Aware Interventions: A Case Study in Chess." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
BibTeX:@article{narayanan2026improving,
title={Improving Human Performance with Value-Aware Interventions: A Case Study in Chess},
author={Saumik Narayanan and Raja Panjwani and Siddhartha Sen and Chien-Ju Ho},
journal={Reinforcement Learning Journal},
volume={7},
pages={},
year={2026}
}