This is the pre-proceedings for the RLC 2026. You may expect minor changes.

Reward-Conditioned Attention: How Reward Design Shapes What Autonomous Driving Agents See

By Mohamed Benabdelouahad, AHMED DJALAL HACINI, Nadir Farhi, and Aissa Boulmerka

Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

Will be presented at the Reinforcement Learning Conference (RLC), Montréal, Quebec, Canada, August 15–17, 2026.


Download:

Keywords: Explainable AI, reinforcement learning, autonomous driving, attention mechanism,

Abstract:

We investigate how reward design shapes the internal attention patterns of reinforcement learning agents trained for autonomous driving. Using three Perceiver-based agents that share identical architectures and training data but differ only in their reward configurations—ranging from basic violation penalties to continuous proximity penalties—we analyze cross-attention allocation across 50 real-world scenarios from the Waymo Open Motion Dataset. A central methodological finding is that naive pooling of timesteps across episodes substantially underestimates the attention–risk relationship; within-episode correlation with Fisher z-transform aggregation is the appropriate statistic and reveals a robustly positive link between collision risk and agent-directed attention. Building on this validated methodology, we demonstrate two reward-conditioned effects: agents trained with navigation rewards allocate up to $2.0\times$ more attention to GPS-path tokens than those trained with additional proximity penalties—and $4.7\times$ more than agents with no navigation incentive—revealing that reward content directly determines which scene elements the encoder prioritizes, and continuous time-to-collision penalties create a learned vigilance prior—elevated resting agent surveillance maintained throughout collision-free phases. In several scenarios, the complete-reward and minimal-reward models exhibit opposite attention–risk correlation directions, demonstrating that reward design can qualitatively reverse attentional strategy rather than merely modulating its magnitude. These results suggest that attention analysis is a practical diagnostic for verifying that a reward function produces the intended representational behaviour in safety-critical RL systems.


Citation Information:

Mohamed Benabdelouahad, AHMED DJALAL HACINI, Nadir Farhi, and Aissa Boulmerka. "Reward-Conditioned Attention: How Reward Design Shapes What Autonomous Driving Agents See." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

BibTeX:
@article{benabdelouahad2026rewardconditioned,
    title={Reward-Conditioned Attention: How Reward Design Shapes What Autonomous Driving Agents See},
    author={Mohamed Benabdelouahad and AHMED DJALAL HACINI and Nadir Farhi and Aissa Boulmerka},
    journal={Reinforcement Learning Journal},
    volume={7},
    pages={},
    year={2026}
}