This is the pre-proceedings for the RLC 2026. You may expect minor changes.

Leveraging Reward Machines for Efficient Multi-Objective Reinforcement Learning

By Panos Aronis, Mehdi Dastani, Roxana Rădulescu, and Giovanni Varricchione

Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

Will be presented at the Reinforcement Learning Conference (RLC), Montréal, Quebec, Canada, August 15–17, 2026.


Download:

Keywords: Multi-Objective Reinforcement Learning, Reward Machines, Reinforcement Learning

Abstract:

Reinforcement Learning (RL) provides a powerful framework for sequential decision-making, typically assuming a single scalar reward function and a Markovian reward structure. However, many real-world problems involve multiple, potentially conflicting objectives and require temporally extended behaviours that fundamentally violate the Markov property. These complexities have separately motivated the study of Multi-Objective Reinforcement Learning (MORL) and Reward Machines (RMs), both of which demand new representations and learning strategies to ensure effective, sample-efficient, and interpretable solutions. To address these challenges, we propose a novel framework, Multi-Objective Reward Machines (MORMs), to handle non-Markovian reward structure and multiple objectives in a unified setting, enabling more principled and sample-efficient learning in complex sequential decision-making tasks.


Citation Information:

Panos Aronis, Mehdi Dastani, Roxana Rădulescu, and Giovanni Varricchione. "Leveraging Reward Machines for Efficient Multi-Objective Reinforcement Learning." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

BibTeX:
@article{aronis2026leveraging,
    title={Leveraging Reward Machines for Efficient Multi-Objective Reinforcement Learning},
    author={Panos Aronis and Mehdi Dastani and Roxana Rădulescu and Giovanni Varricchione},
    journal={Reinforcement Learning Journal},
    volume={7},
    pages={},
    year={2026}
}