This is the pre-proceedings for the RLC 2026. You may expect minor changes.
Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.
Keywords: temporal credit assignment, counterfactual Shapley values, causal inference
The Credit Assignment Problem (CAP) is fundamental to developing efficient and explainable Reinforcement Learning (RL) agents. Existing frameworks, whether relying on temporal contiguity or hindsight-conditioned reweighting of rewards, frequently fail to distribute credit properly between an agent's policy (skill) and environmental stochasticity (luck). A principled approach to the CAP must isolate the causes of observed rewards from spurious correlated features and environmental randomness. We introduce Counterfactual Shapley Credit Assignment, a novel causal credit assignment framework based on the counterfactual Shapley value ($\phi$-value). By redistributing rewards, $\phi$-values enhance credit assignment across three critical dimensions: high stochasticity, sparse causality, and delayed rewards, all while theoretically preserving the original optimal policy. We derive a consistent estimator that computes each $\phi$-value in amortized constant time complexity, enabling a new class of policy gradient methods, $\phi$-PPO. Empirical results demonstrate that $\phi$-values align precisely with the ground truth causes of task rewards. Furthermore, we show its superior sample efficiency in challenging environments where prior state-of-the-art methods fail to converge.
Mingxuan Li, Kai-Zhan Lee, and Elias Bareinboim. "Counterfactual Shapley Credit Assignment." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
BibTeX:@article{li2026counterfactual,
title={Counterfactual Shapley Credit Assignment},
author={Mingxuan Li and Kai-Zhan Lee and Elias Bareinboim},
journal={Reinforcement Learning Journal},
volume={7},
pages={},
year={2026}
}