This is the pre-proceedings for the RLC 2026. You may expect minor changes.
Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.
Keywords: natural policy gradient, Bayesian network, Markov potential games
We study natural policy gradient (NPG) with Bayesian network (BN) policies in Markov potential games (MPGs) under softmax parameterization. BN policies represent correlated joint policies via a directed acyclic graph, generalizing the independent (product) policies used in prior work. We prove that BN-NPG converges to near-Nash policies at an $O(1/K)$ rate averaged over $K$ iterations for any fixed DAG, and that for fully correlated BN policies, every near-Nash policy is also near-optimal, recovering the global optimality guarantee of single-agent NPG. Unlike product policies, BN policies condition each agent's action on parent actions, requiring sufficient exploration over all parent-action profiles. We impose this as an assumption in the unregularized setting, and then remove it by designing a log-barrier regularizer, weighted by the frequency of encountering each parent-action profile under the current policy, that provably maintains exploration throughout training. We further extend these results to partially observable settings with factored state spaces, where each agent observes only its local state.
Dingyang Chen, Zhenyu Zhang, Yuan Ling, and Qi Zhang. "Natural Policy Gradient for Bayesian Network Policies in Markov Potential Games." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
BibTeX:@article{chen2026natural,
title={Natural Policy Gradient for Bayesian Network Policies in Markov Potential Games},
author={Dingyang Chen and Zhenyu Zhang and Yuan Ling and Qi Zhang},
journal={Reinforcement Learning Journal},
volume={7},
pages={},
year={2026}
}