This is the pre-proceedings for the RLC 2026. You may expect minor changes.

Overcoming Valid Action Suppression in Unmasked Policy Gradient Algorithms

By Renos Zabounidis, Roy Siegelmann, Mohamad Qadri, Woojun Kim, Simon Stepputtis, and Katia P. Sycara

Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.


Download:

Keywords: Action masking, policy gradient, valid action suppression, feasibility classification

Abstract:

In reinforcement learning environments with state-dependent action validity, action masking consistently outperforms penalty-based handling of invalid actions, yet existing theory only shows that masking preserves the policy gradient theorem. We identify a distinct failure mode of unmasked training: it systematically suppresses valid actions at states the agent has not yet visited. This occurs because gradients pushing down invalid actions at visited states propagate through shared network parameters to unvisited states where those actions are valid. We prove that for softmax policies with shared features, when an action is invalid at visited states but valid at an unvisited state $s^\ast$, the probability $\pi(a \mid s^*)$ decays exponentially in the number of gradient steps, through parameter sharing and a zero-sum identity of softmax logit updates. Entropy regularization is reactive rather than preventive: it provides no direct update at unvisited states, and opposes suppression only after the state is visited and the action probability has become sufficiently small. We validate empirically that deep networks exhibit the feature alignment condition required for suppression, and experiments on Craftax, Craftax-Classic, and MiniHack confirm the predicted exponential suppression and show that training with a feasibility-classification loss improves performance over oracle masking alone, while also enabling deployment without oracle masks.


Citation Information:

Renos Zabounidis, Roy Siegelmann, Mohamad Qadri, Woojun Kim, Simon Stepputtis, and Katia P. Sycara. "Overcoming Valid Action Suppression in Unmasked Policy Gradient Algorithms." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

BibTeX:
@article{zabounidis2026overcoming,
    title={Overcoming Valid Action Suppression in Unmasked Policy Gradient Algorithms},
    author={Renos Zabounidis and Roy Siegelmann and Mohamad Qadri and Woojun Kim and Simon Stepputtis and Katia P. Sycara},
    journal={Reinforcement Learning Journal},
    volume={7},
    pages={},
    year={2026}
}