This is the pre-proceedings for the RLC 2026. You may expect minor changes.

Escaping Offline Pessimism: Vector-Field Reward Shaping for Safe Frontier Exploration

By Amirhossein Roknilamouki, Arnob Ghosh, Eylem Ekici, and Ness Shroff

Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.


Download:

Keywords: offline reinforcement learning, safe exploration, reward shaping, fixed-policy

Abstract:

Offline reinforcement learning often produces conservative policies whose pessimism limits online exploration and data collection. Inspired by safe reinforcement learning, we target the boundary of regions well covered by offline data and reliably modeled by the simulator, where samples are informative but uncertainty remains moderate. However, naively rewarding this boundary-seeking behavior during offline training can lead to a degenerate parking behavior at deployment, where the fixed policy stops once it reaches the frontier. To address this issue, we propose a novel vector-field reward-shaping paradigm designed to induce continuous boundary exploration for fixed policies at deployment. Assuming access to a differentiable uncertainty oracle, our reward combines two complementary components: a gradient-alignment term that attracts the agent toward a target uncertainty level, and a rotational-flow term that promotes motion along the local tangent plane of the uncertainty manifold. Our average-reward analysis shows that the policy jointly optimizes the primary MDP reward and exploratory motion, with the trade-off characterized through its stationary state-visitation distribution. In a 2D continuous-navigation task with an analytically specified uncertainty oracle, integrating the reward with Soft Actor-Critic induces sustained boundary exploration while retaining goal-directed behavior.


Citation Information:

Amirhossein Roknilamouki, Arnob Ghosh, Eylem Ekici, and Ness Shroff. "Escaping Offline Pessimism: Vector-Field Reward Shaping for Safe Frontier Exploration." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

BibTeX:
@article{roknilamouki2026escaping,
    title={Escaping Offline Pessimism: Vector-Field Reward Shaping for Safe Frontier Exploration},
    author={Amirhossein Roknilamouki and Arnob Ghosh and Eylem Ekici and Ness Shroff},
    journal={Reinforcement Learning Journal},
    volume={7},
    pages={},
    year={2026}
}