This is the pre-proceedings for the RLC 2026. You may expect minor changes.

The Yokai Learning Environment: Tracking Beliefs Over Space and Time

By Constantin Ruhdorfer, Matteo Bortoletto, Johannes Forkel, Jakob Nicolaus Foerster, and Andreas Bulling

Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

Will be presented at the Reinforcement Learning Conference (RLC), Montréal, Quebec, Canada, August 15–17, 2026.


Download:

Keywords: Multi-Agent Reinforcement Learning, Zero-Shot Coordination, Computational Theory

Abstract:

The ability to cooperate with unknown partners is a central challenge in cooperative AI and widely studied in the form of zero-shot coordination (ZSC), which evaluates an algorithm by measuring the performance of independently trained agents when paired. The Hanabi Learning Environment (HLE) has become the dominant benchmark for ZSC, but recent work has achieved near-perfect inter-seed cross-play performance, limiting its ability to track algorithmic progress. We introduce the Yokai Learning Environment (YLE) for ZSC - an open-source multi-agent RL benchmark in which effective collaboration among unknown agents requires building common ground by tracking and updating beliefs over moving cards, reasoning under ambiguous hints, and deciding when to terminate the game based on inferred shared knowledge - features absent in the HLE, where beliefs are tied to few hand slots and hints are truthful by rule. We evaluate three leading ZSC methods: High-Entropy IPPO, Other-Play, and Off-Belief Learning, which achieve near-perfect inter-seed cross-play in the HLE, and show that in the YLE they exhibit persistent SP–XP gaps, degraded early-ending calibration, and weaker belief representations in cross-play, indicating failure to maintain consistent internal models with unseen partners. Methods that perform best in the HLE do not perform best in the YLE, indicating that progress measured on a single benchmark may not generalise. Together, these results establish YLE as a challenging new ZSC benchmark.


Citation Information:

Constantin Ruhdorfer, Matteo Bortoletto, Johannes Forkel, Jakob Nicolaus Foerster, and Andreas Bulling. "The Yokai Learning Environment: Tracking Beliefs Over Space and Time." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

BibTeX:
@article{ruhdorfer2026the,
    title={The Yokai Learning Environment: Tracking Beliefs Over Space and Time},
    author={Constantin Ruhdorfer and Matteo Bortoletto and Johannes Forkel and Jakob Nicolaus Foerster and Andreas Bulling},
    journal={Reinforcement Learning Journal},
    volume={7},
    pages={},
    year={2026}
}