This is the pre-proceedings for the RLC 2026. You may expect minor changes.
Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.
Keywords: Human-AI Interaction, Trust Signaling, Preemptive Commitment, Legible Avoidance.
In human-AI interaction, a low-trust environment could emerge when users are uncertain about the agent's capabilities, intentions, or reliability. This may hinder not only effective collaboration but also the adoption of extremely beneficial systems. In this study, we present a general framework to model such environments and a strategy to improve user trust, namely, through signaling behavior. Specifically, we propose two signaling techniques for an AI agent in low-trust environments. The first strategy is called preemptive commitment, in which the agent attempts to drive itself into a state from which it cannot violate any of the normative contracts. The second strategy is called legible avoidance, which involves the agent conveying that it is not trying to violate a contract in the normative contracts by moving away from the object of interest. We evaluated the performance of these strategies across standard MDP planning benchmarks and conducted a user study to investigate how users perceive preemptive commitment and legible avoidance signaling techniques, particularly their trust level when the agent demonstrates these techniques.
Septia Rani, Turgay Caglar, and Sarath Sreedharan. "Planning for Signaling in Low-Trust Environments." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.
BibTeX:@article{rani2026planning,
title={Planning for Signaling in Low-Trust Environments},
author={Septia Rani and Turgay Caglar and Sarath Sreedharan},
journal={Reinforcement Learning Journal},
volume={7},
pages={},
year={2026}
}