This is the pre-proceedings for the RLC 2026. You may expect minor changes.

Counterfactual Shapley Credit Assignment

By Mingxuan Li, Kai-Zhan Lee, and Elias Bareinboim

Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.


Download:

Keywords: temporal credit assignment, counterfactual Shapley values, causal inference

Abstract:

The Credit Assignment Problem (CAP) is fundamental to developing efficient and explainable Reinforcement Learning (RL) agents. Existing frameworks, whether relying on temporal contiguity or hindsight-conditioned reweighting of rewards, frequently fail to distribute credit properly between an agent's policy (skill) and environmental stochasticity (luck). A principled approach to the CAP must isolate the causes of observed rewards from spurious correlated features and environmental randomness. We introduce Counterfactual Shapley Credit Assignment, a novel causal credit assignment framework based on the counterfactual Shapley value ($\phi$-value). By redistributing rewards, $\phi$-values enhance credit assignment across three critical dimensions: high stochasticity, sparse causality, and delayed rewards, all while theoretically preserving the original optimal policy. We derive a consistent estimator that computes each $\phi$-value in amortized constant time complexity, enabling a new class of policy gradient methods, $\phi$-PPO. Empirical results demonstrate that $\phi$-values align precisely with the ground truth causes of task rewards. Furthermore, we show its superior sample efficiency in challenging environments where prior state-of-the-art methods fail to converge.


Citation Information:

Mingxuan Li, Kai-Zhan Lee, and Elias Bareinboim. "Counterfactual Shapley Credit Assignment." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

BibTeX:
@article{li2026counterfactual,
    title={Counterfactual Shapley Credit Assignment},
    author={Mingxuan Li and Kai-Zhan Lee and Elias Bareinboim},
    journal={Reinforcement Learning Journal},
    volume={7},
    pages={},
    year={2026}
}