RLJ 2026: Volume 7
This is the pre-proceedings for the RLC 2026. Below are links to individual papers.
- Overcoming Valid Action Suppression in Unmasked Policy Gradient Algorithms, by Renos Zabounidis, Roy Siegelmann, Mohamad Qadri, Woojun Kim, Simon Stepputtis, and Katia P. Sycara.
- Centralized Adaptive Sampling for Reliable Co-training of Independent Multi-Agent Policies, by Nicholas E. Corrado and Josiah P. Hanna.
- Yes, Q-learning Helps Offline In-Context RL, by Denis Tarasov, Alexander Nikulin, Ilya Zisman, Albina Klepach, Andrei Polubarov, Lyubaykin Nikita, Alexander Derevyagin, Igor Kiselev, and Vladislav Kurenkov.
- On the Variance of Temporal Difference Learning and its Reduction Using Control Variates, by Hsiao-Ru Pan and Bernhard Schölkopf.
- Deep Reinforcement Learning for Spacecraft Attitude Control During Atmospheric Re-Entry, by Alexander Fabisch, Melvin Laux, Mariela De Lucas Alvarez, Edoardo Caroselli, and Julian Theis.
- Towards Formalizing Reinforcement Learning Theory: A Robbins-Siegmund Approach, by Shangtong Zhang.
- Statistical Inference for Policy Evaluation with Temporal Difference Learning, by Weichen Wu, Gen Li, Yuting Wei, and Alessandro Rinaldo.
- PPO+: Enhancing proximal policy optimization, by Mahdi Kallel, Jose-Luis Holgado-Alvarez, Samuele Tosatto, and Carlo D'Eramo.
- Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control, by Rohit Kumar Salla, Manoj Saravanan, and Simon Stepputtis.
- Limits of reinforcement learning for decision trees in Markov decision processes, by Hector Kohler, Riad Akrour, and Philippe Preux.
- Direct Advantage Estimation for Scalable and Sample-efficient Deep Reinforcement Learning, by Hsiao-Ru Pan and Bernhard Schölkopf.
- From Demonstrations to Rewards: Test-Time Prompt Optimization for VLM Reward Models, by Christian Gumbsch, Leonardo Barcellona, Lennard Schuenemann, Platon Karageorgis, Andrii Zadaianchuk, Zehao Wang, Sergey Zakharov, Fabien Despinoy, Rahaf Aljundi, and Stratis Gavves.
- Biased Dreams: Limitations to Epistemic Uncertainty Quantification in Latent Dynamics Models, by Julia Berger, Bernd Frauenknecht, Sebastian Trimpe, and Bastian Leibe.
- Learning in Low-Dimensional Subspaces: Orthogonal Bottlenecks for Reinforcement Learning, by Aleksandar Todorov and Matthia Sabatelli.
- CODA: Coordination via On-Policy Diffusion for Multi-Agent Offline Reinforcement Learning, by Marcel Hedman, Kale-ab Tessera, Juan Claude Formanek, Anya Sims, Riccardo Zamboni, Trevor McInroe, John Torr, and Elliot Fosong.
- Repetition as Reinforcement: Enhancing Sample Efficiency via Instant Episode Repetition in Reinforcement Learning, by Hoda Yamani, Yuning Xing, Koen van Rijnsoever, Bruce A. MacDonald, and Henry Williams.
- Risk-Aware General-Utility Markov Decision Processes, by Pedro Pinto Santos, Fábio Vital, Alberto Sardinha, and Francisco S. Melo.
- The Yokai Learning Environment: Tracking Beliefs Over Space and Time, by Constantin Ruhdorfer, Matteo Bortoletto, Johannes Forkel, Jakob Nicolaus Foerster, and Andreas Bulling.
- Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments, by Michael Beukman, Khimya Khetarpal, Zeyu Zheng, Will Dabney, Jakob Nicolaus Foerster, Michael D Dennis, and Clare Lyle.
- Stable Planning through Aligned Representations in Model-Based Reinforcement Learning, by Misagh Soltani and Forest Agostinelli.
- Delightful Policy Gradient, by Ian Osband.
- Beyond Local Views: Global State Inference with Diffusion Models for Cooperative MARL, by Zhiwei Xu, Hangyu Mao, ZHANG NIANMIN, Shengtao Zhang, Xin Xin, Pengjie Ren, Dapeng Li, Bin Zhang, Guoliang Fan, Zhumin Chen, Changwei Wang, and Jiangjin Yin.
- The Cell Must Go On: Agar.io for Continual Reinforcement Learning, by Mohamed Ayman Mohamed, Kateryna Nekhomiazh, Vedant Vyas, Marcos Menon Jose, Andrew Patterson, and Marlos C. Machado.
- Momba: Network Modernization Improves Multi-Objective Reinforcement Learning, by Adam Štafa, Santeri Heiskanen, Petr Novotný, and Joni Pajarinen.
- ICPL: Few-shot In-context Preference Learning via LLMs, by Chao Yu, Qixin Tan, Hong Lu, Jiaxuan Gao, Xinting Yang, Yu Wang, Yi Wu, and Eugene Vinitsky.
- Natural Policy Gradient for Bayesian Network Policies in Markov Potential Games, by Dingyang Chen, Zhenyu Zhang, Yuan Ling, and Qi Zhang.
- Gradient Iterated Temporal-Difference Learning, by Théo Vincent, Kevin Gerhardt, Yogesh Tripathi, Habib Maraqten, Adam White, Martha White, Jan Peters, and Carlo D'Eramo.
- Confidence Intervals for the Interquartile Mean, by Alexandra Burushkina and Philip S. Thomas.
- Temporally Extended Mixture-of-Experts Models, by Zeyu Shen and Peter Henderson.
- Synthetic Monitoring Environments for Reinforcement Learning, by Leonard S. Pleiss, Carolin Schmidt, and Maximilian Schiffer.
- Memory Retention Is Not Enough to Master Memory Tasks in Reinforcement Learning, by Oleg Shchendrigin, Egor Cherepanov, Alexey Kovalev, and Aleksandr Panov.
- PGTG: Procedurally Generated Grid-Based Traffic Gym, by Joshua Meyer, Felix Maurice Kuntz, Verena Wolf, Jörg Hoffmann, and Timo P. Gros.
- Extending Differential Temporal Difference Methods for Episodic Problems, by Kris De Asis, Mohamed Elsayed, and Jiamin He.
- Dense and Diverse Goal Coverage in Multi Goal Reinforcement Learning, by Sagalpreet Singh, Rishi Saket, and Aravindan Raghuveer.
- Credit Assignment and Focused Exploration for Sparse-reward Multi-agent Deep Reinforcement Learning, by Shuai Han, Mehdi Dastani, and Shihan Wang.
- Multi-Agent Reinforcement Learning with Reward Machines for Mixed Cooperative-Competitive Environments, by Sriram Ganapathi Subramanian, Toryn Q. Klassen, and Sheila A. McIlraith.
- StaQ: a Finite Memory Approach to Discrete Action Policy Mirror Descent, by Alex Davey, Alena Shilova, Brahim Driss, and Riad Akrour.
- Bi-Level Reinforcement Learning Pathway for Sim-to-Real Optimality, by Akhil S Anand, Shambhuraj Sawant, Paavo Parmas, Jasper Hoffmann, Dirk Reinhardt, and Sebastien Gros.
- Forager: a lightweight testbed for continual learning with partial observability in RL, by Steven Tang, Xinze Xiong, Anna Hakhverdyan, Andrew Patterson, Jacob Adkins, Jiamin He, Esraa Elelimy, Parham Mohammad Panahi, Martha White, and Adam White.
- Near-Optimal Reinforcement Learning for Linear Distributionally Robust Markov Decision Processes, by Zhishuai Liu, Weixin Wang, and Pan Xu.
- Revisiting FTA: A Sparse One-to-Many Activation for Reinforcement Learning, by Tyler Lazar, Matthew Vandergrift, Martha White, and Adam White.
- Short-Term-to-Long-Term Memory Transfer for Knowledge Graphs under Partial Observability, by Taewoon Kim, Vincent Francois-Lavet, and Michael Cochez.
- On the Sample Complexity of Discounted Reinforcement Learning with Optimized Certainty Equivalents, by Oliver Mortensen and M. Sadegh Talebi.
- Randomized Exploration for Linear Bandits via Absolute Perturbations, by Toshinori Kitamura, Shuai Liu, and Csaba Szepesvari.
- Counterfactual Shapley Credit Assignment, by Mingxuan Li, Kai-Zhan Lee, and Elias Bareinboim.
- Ludax: A GPU-Accelerated Description Language for Board Games, by Graham Todd, Alexander George Padula, Dennis J. N. J. Soemers, Sam Earle, and Julian Togelius.
- Scalable Causal Imitation Learning, by Eylam Tagor, Mingxuan Li, and Elias Bareinboim.
- Intrinsic Closed-Loop Practical Asymptotic Stability in Discrete-Time Reinforcement Learning, by Jan de Priester and Ricardo Sanfelice.
- Toward Agents That Reason About Their Computation, by Adrian Orenstein, Jessica Chen, Gwyneth Anne Delos Santos, Bayley Sapara, and Michael Bowling.
- Learning the Supports for Categorical Critic in Reinforcement Learning, by Jen-Yen Chang, Takayuki Osa, and Tatsuya Harada.
- Modification-Considering Value Learning for Reward Hacking Mitigation in RL, by Evgenii Opryshko, Umangi Jain, and Igor Gilitschenski.
- Using Common Random Numbers for Simulation-based Planning with Rollouts, by Sandarbh Yadav, Frederic J Maliakkal, Harshad Khadilkar, and Shivaram Kalyanakrishnan.
- Trajectory First: A Curriculum for Discovering Diverse Policies, by Cornelius V. Braun, Sayantan Auddy, and Marc Toussaint.
- Rethinking the Suitability of RL Algorithms Under Practical Transfer Constraints, by Hany Hamed, Abhishek Naik, Colin Bellinger, and A. Rupam Mahmood.
- Learning Multi-Agent Communication Protocol: Study on Information Entropy Efficiency in MARL, by Xinren Zhang, Zixin Zhong, and Jiadong Yu.
- Finite-Time Analysis of the Natural Policy Gradient in Finite-Horizon Markov Decision Processes, by Asha Barua and Sajad Khodadadian.
- V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control, by Donghu Kim, Youngdo Lee, Hojoon Lee, Johan Obando-Ceron, Byungkun Lee, Aaron Courville, Pablo Samuel Castro, Jaegul Choo, and Clare Lyle.
- Coordination Graphs for Constrained Multi-Agent Reinforcement Learning, by Santiago Amaya-Corredor, Miguel Calvo-Fullana, and Anders Jonsson.
- Maximum-Entropy Exploration with Future State-Action Visitation Measures, by Adrien Bolland, Gaspard Lambrechts, and Damien Ernst.
- Fully Offline Reinforcement Learning, by Mattie Fellows, Clarisse Wibault, Uljad Berdica, Johannes Forkel, Michael A Osborne, and Jakob Nicolaus Foerster.
- Training on Irrelevant States Implies Data Augmentation: Generalization in Contextual MDPs, by Max Weltevrede, Caroline Horsch, Matthijs T. J. Spaan, and Wendelin Boehmer.
- Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function, by Qining Zhang and Lei Ying.
- Conformal Preemption of Failures in Sequential Decision-Making Agents, by Garrett Ethan Katz, Adebayo Braimah, Qinru Qiu, and Simon Khan.
- PB²: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning, by Brahim Driss, Alex Davey, and Riad Akrour.
- Dynamics Models for Offline Hyperparameter Selection in Real-World RL, by Jordan Coblin, Han Wang, Martha White, and Adam White.
- An Agent-Centric Dynamical Systems Perspective on Multi-Agent Reinforcement Learning, by James Rudd-Jones, Maria Perez-Ortiz, and Mirco Musolesi.
- Let it Cook: Learning to Wait in Sequential Decision Making, by Christopher Watson, Arjun Krishna, Dinesh Jayaraman, and Rajeev Alur.
- Provable Distributional Value Iteration under Partial Observability, by Larry Preuett, Qiuyi Zhang, and Muhammad Aurangzeb Ahmad.
- Reward-Conditioned Attention: How Reward Design Shapes What Autonomous Driving Agents See, by Mohamed Benabdelouahad, AHMED DJALAL HACINI, Nadir Farhi, and Aissa Boulmerka.
- Prediction-Based Markov Violation Scores for Detecting Non-Markovian Observations in Reinforcement Learning, by Naveen Mysore.
- Distributionally Robust Self Paced Curriculum Reinforcement Learning, by Anirudh Satheesh, Keenan Powell, and Vaneet Aggarwal.
- Gaussian Process Aggregation for Root-Parallel Monte Carlo Tree Search with Continuous Actions, by Junlin Xiao, Victor-Alexandru Darvariu, Bruno Lacerda, and Nick Hawes.
- ContrastSpanner: Learning Low-Rank Causal Contrasts to Alleviate Power-Set and Eluder Barriers, by Alec Koppel and Laixi Shi.
- DART: Dual Adaptive Residual Tracking for Low-Bias Advantage Estimation and Credit Assignment in AI Agents, by Shahrad Mohammadzadeh, Amir-massoud Farahmand, Reihaneh Rabbany, and Doina Precup.
- Goal-Oriented Reinforcement Learning for Stochastic Shortest Paths with Dead-Ends, by Gustavo De Mari Pereira and Leliane N. de Barros.
- ASALT: Adaptive State Alignment for Lateral Transfer in Multi-agent Reinforcement Learning, by Anurag Akula, Satheesh K Perepu, Abhishek Sarkar, and Kaushik Dey.
- Learning with Coupled Uncertainty, by Waqar Mirza, Aldo Pacchiano, and Eric Mazumdar.
- NutriRL: A Benchmark for Nutritional Regulation under Delayed State Transitions, by Aniket Khan, Charitha Palika, and V.Srinivasa Chakravarthy.
- Representation Regularization in Distributional Reinforcement Learning, by André Inge, Jonas Nordqvist, Björn Lindenberg, and Karl-Olof Lindahl.
- Strategically-Linked Decisions in Long-Term Planning and Reinforcement Learning, by Alihan Hüyük, Jonas B Raedler, Leo Benac, and Finale Doshi-Velez.
- Reward Design Agent for Reinforcement Learning, by Hojoon Lee, Ajay Subramanian, Ben Abbatematteo, Vijay Veerabadran, Pedro Matias, Karl Ridgeway, and Nitin Kamra.
- Safe Flow Q-Learning: Offline Safe Reinforcement Learning with Reachability-Based Flow Policies, by Mumuksh Tayal, Manan Tayal, and Ravi Prakash.
- A Causality-Inspired Spatial-Temporal Return Decomposition Approach for Multi-Agent Reinforcement Learning, by Yudi Zhang, Yali Du, Biwei Huang, Mykola Pechenizkiy, and Meng Fang.
- Simple Recipe Works: Vision-Language-Action Models are Natural Continual Learners with Reinforcement Learning, by Jiaheng Hu, Jay Shim, Chen Tang, Yoonchang Sung, Bo Liu, Peter Stone, and Roberto Martín-Martín.
- Supervised Reward Inference, by Will Schwarzer, Jordan Jack Schneider, Philip S. Thomas, and Scott Niekum.
- Minimal Ingredients for Reward Assignment from Expert Demonstrations, by Zixuan Dong, Yumi Omori, and Keith W. Ross.
- Solvable models of learning to pursue a moving target, by John J. Vastola and Kanaka Rajan.
- Shielded Controller Units for RL with Operational Constraints Applied to Remote Microgrids, by Hadi Nekoei, Alexandre Blondin Massé, Rachid Hassani, Sarath Chandar, and Vincent Mai.
- Adaptive Critic Shaping for Reinforcement Learning with Temporal Logic Constraint, by Duo XU.
- Weight a moment: risk-aware exploration with Bayes, by Karim Zaghw, Peter Dayan, and Georgy Antonov.
- When Do We Need LLMs? A Diagnostic for Language-Driven Bandits, by Uljad Berdica, Fernando Acero, Anton Ipsen, Parisa Zehtabi, Michael Cashmore, and Manuela Veloso.
- Uncertainty-Aware Predictive Safety Filters for Probabilistic Neural Network Dynamics, by Bernd Frauenknecht, Lukas Kesper, Daniel Mayfrank, Henrik Hose, and Sebastian Trimpe.
- Discovering Reinforcement Learning Interfaces with Large Language Models, by Akshat Singh Jaswal, Ashish Baghel, and Paras Chopra.
- When to Plan: Learning to Select Between Reactive Control and Deliberative Planning, by Adam Labiosa and Josiah P. Hanna.
- Maximum Entropy Behavior Exploration for Sim2Real Zero-Shot Reinforcement Learning, by Jiajun Hu, Núria Armengol Urpí, Jin Cheng, and Stelian Coros.
- An Unreasonably Simple Approach to Safe RL, by Geraud Nangue Tasse, Mark Nemecek, Tamlin Love, Steven James, and Benjamin Rosman.
- SCoUT: Scalable Communication via Utility-Guided Temporal Grouping in Multi-Agent Reinforcement Learning, by Manav Vora, Gokul Puthumanaillam, Hiroyasu Tsukamoto, and Melkior Ornik.
- SegDAC: Visual Generalization in Reinforcement Learning via Dynamic Object Tokens, by Alexandre Brown and Glen Berseth.
- Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification, by Haoyang Hong, Zichen Wang, Quanquan Gu, and Huazheng Wang.
- Conservative Value Priors: A Bayesian Path to Offline Reinforcement Learning, by Filippo Valdettaro, Yingzhen Li, and Aldo A. Faisal.
- Rationalizing Boltzmann Rationality: An Axiomatic Characterization of Entropy-Regularized Policies, by Silviu Pitis.
- Physical Atari: A Robust and Accessible Platform for Real-time Reinforcement Learning on Robots, by Khurram Javed, Joseph Varughese Modayil, Gloria Kennickell, Richard S Sutton, and John Carmack.
- Building2Building: A Large Scale Benchmark for Generalizable Real-World Reinforcement Learning, by Vincent Taboga, Justin Veilleux, Doseok Jang, Anushree Rankawat, and Pierre-Luc Bacon.
- Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning, by Ali Larian, Qian Lin, Chang Zong Wu, and Daniel S. Brown.
- Approximate Next Policy Sampling: Replacing Conservative Target Policy Updates in Deep RL, by Dillon Sandhu and Ronald Parr.
- Hierarchical Behaviour Spaces, by Michael Matthews, Pierluca D'Oro, Anssi Kanervisto, Scott Fujimoto, Jakob Nicolaus Foerster, and Mikael Henaff.
- Planning for Signaling in Low-Trust Environments, by Septia Rani, Turgay Caglar, and Sarath Sreedharan.
- Towards Affordable Energy: A Gymnasium Environment for Electric Utility Demand-Response Program, by Jose Efraim Aguilar Escamilla, Lingdong Zhou, Xiangqi Zhu, and Huazheng Wang.
- FlowRL: A Taxonomy and Modular Framework for Reinforcement Learning with Diffusion Policies, by Chenxiao Gao, Edward Chen, Tianyi Chen, and Bo Dai.
- Assistax: A Multi-Agent Hardware-Accelerated Reinforcement Learning Benchmark for Assistive Robotics, by Leonard Hinckeldey, Elliot Fosong, Rimvydas Rubavicius, Elle Miller, Trevor McInroe, Fan Zhang, Patricia Wollstadt, Stefano V. Albrecht, and Subramanian Ramamoorthy.
- A Value-Based Approach to Maximum Entropy Exploration, by Jacob Adamczyk, Adam Kamoski, and Rahul V Kulkarni.
- Optimal Regret for Policy Optimization in Average Reward MDPs Without Mixing, by William Powell, Jeongyeol Kwon, Qiaomin Xie, and Hanbaek Lyu.
- A Simple Baseline for Learning Approximate State Abstractions in Factored State Spaces, by Anshuman Senapati and Josiah P. Hanna.
- Strategically Robust Multi-Agent Reinforcement Learning with Linear Function Approximation, by Jake Gonzales, Max Horwitz, Eric Mazumdar, and Lillian J. Ratliff.
- From Pixels to Factors: Learning Independently Controllable State Variables for Reinforcement Learning, by Rafael Rodriguez-Sanchez, Cameron Allen, and George Konidaris.
- Escaping Offline Pessimism: Vector-Field Reward Shaping for Safe Frontier Exploration, by Amirhossein Roknilamouki, Arnob Ghosh, Eylem Ekici, and Ness Shroff.
- Decentralized Asymmetric DQN: Decentralization without Factorization in Multi-Agent Reinforcement Learning, by Rupali Bhati, Anurag Kadkol, Andrea Baisero, and Christopher Amato.
- When Can Pure Exploitation Succeed in Linear RL? Decoys and Self-Identifiability for Greedy LSVI, by Manoj Saravanan and Rohit Kumar Salla.
- Inference-Time Policy Alignment for Fair Reinforcement Learning, by Umer Siddique, Peilang Li, Conor Wallace, and Yongcan Cao.
- Offline RL with Hierarchical Action Chunking, by Ahad Jawaid.
- Fixing Incomplete Value Function Decomposition for Multi-Agent Reinforcement Learning, by Andrea Baisero, Rupali Bhati, Shuo Liu, Aathira Sunil Pillai, and Christopher Amato.
- Endpoint Replay: Compressing the Recency Buffer in Deep Reinforcement Learning, by Parham Mohammad Panahi, Armin Ashrafi, Haoyu Du, Andrew Patterson, Martha White, and Adam White.
- Improving Human Performance with Value-Aware Interventions: A Case Study in Chess, by Saumik Narayanan, Raja Panjwani, Siddhartha Sen, and Chien-Ju Ho.
- Learning World Value Functions with Successor Representation and Vision Models, by Sergio Frasco, Devon Jarvis, and Geraud Nangue Tasse.
- Best-of-Both-Worlds Multi-Dueling Bandits: Unified Algorithms for Stochastic and Adversarial Preferences under Condorcet and Borda Objectives, by S Akash, Pratik Gajane, and Jawar Singh.
- The Open Ant: A Robot Platform for Reinforcement Learning Research, by Elena Sorina Lupu, Patrick Spieler, Khurram Javed, Kris De Asis, John D Martin, Martha Steenstrup, and Joseph Varughese Modayil.
- Offline-to-Online Learning in Linear Bandits, by Kushagra Chandak, Toshinori Kitamura, and Xiaoqi Tan.
- Annealed Softmax Greedy in Many-Armed Bayesian Bandits, by William Overman and Mohsen Bayati.
- Collaborative Learning under Strategic Behavior: Mechanisms for Eliciting Feedback in Principal-Agent Bandit Games, by Ramakrishnan Krishnamurthy, Arpit Agarwal, Lakshmi Subramanian, and Maximilian Nickel.
- Learning Communication Skills in Multi-task Multi-agent Deep Reinforcement Learning, by Changxi Zhu, Mehdi Dastani, and Shihan Wang.
- Revisiting Adam for Streaming Reinforcement Learning, by Florin Gogianu, Luțu Adrian-Cătălin, and Razvan Pascanu.
- Generalize and Guide: Decomposing Rewards for Few-Shot Inverse Reinforcement Learning, by Ziyi Liu and Grace Zhang.
- Improving Reward-Based Hindsight Credit Assignment, by Aditya A. Ramesh, Jiamin He, Jürgen Schmidhuber, and Martha White.
- Bandits for Efficient Experimentation: Adapting to Control Group, Preferences, and Context Drifts, by Udvas Das, Waris Radji, Debabrota Basu, and Odalric-Ambrym Maillard.
- Unsupervised Behavioral Compression: Learning Low-Dimensional Policy Manifolds through State-Occupancy Matching, by Andrea Fraschini, Davide Tenedini, Riccardo Zamboni, Mirco Mutti, and Marcello Restelli.
- Q-Based Variational Inverse Reinforcement Learning, by Ondrej Bajgar, Peter Tisnikar, Konstantinos Gatsis, Alessandro Abate, and Michael A Osborne.
- ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning, by Daiki E. Matsunaga, Junho Na, Tri Wahyu Guntara, Scott Sanner, Pascal Poupart, Jongmin Lee, and Kee-Eung Kim.
- Cohering Reinforcement Learning, by Anna Harutyunyan, Will Dabney, and Doina Precup.
- Generalization in Monitored Markov Decision Processes (Mon-MDPs), by Montaser Mohammedalamen and Michael Bowling.
- Dynamic Object Masks as Goal Representations for Visual Goal-Conditioned Reinforcement Learning, by Fahim Shahriar, Cheryl Wang, Seyed Alireza Azimi, Gautham Vasan, Hany Hamed, Abhishek Naik, A. Rupam Mahmood, and Colin Bellinger.
- Gated Q-learning: Add Off-Policy Bias to Taste, by Brett Daley.
- The challenge of hidden gifts in multi-agent reinforcement learning, by Dane Malenfant and Blake Aaron Richards.
- Leveraging Reward Machines for Efficient Multi-Objective Reinforcement Learning, by Panos Aronis, Mehdi Dastani, Roxana Rădulescu, and Giovanni Varricchione.
- A-3PO: Accelerating Asynchronous LLM Training with Staleness-aware Proximal Policy Approximation, by Xiaocan Li, Shiliang Wu, and Zheng Shen.
- Skill-based Safe Reinforcement Learning with Risk Planning, by Hanping Zhang and Yuhong Guo.
- Design Principles for Tabular Multi-Policy MORL in Infinite Horizons, by Marcelo d'Almeida and Daniel Mosse.
- Discovering High Quality Chess Puzzles with Offline Reinforcement Learning, by Allen Nie, Anirudhan Badrinath, Nicholas Tomlin, Timothy Dai, Carissa Yip, Rose E Wang, Emma Brunskill, and Christopher J Piech.
- Human-Like Goalkeeping in a Realistic Football Simulation: a Sample-Efficient Reinforcement Learning Approach, by Alessandro Sestini, Joakim Bergdahl, Jean-Philippe Barrette-LaPierre, Florian Fuchs, Brady Chen, Fabio Zinno, Michael D Jones, and Linus Gisslén.
- Grounding LTL Tasks in Sub-Symbolic RL Environments for Zero-Shot Generalization, by Matteo Pannacci, Andrea Fanti, Elena Umili, and Roberto Capobianco.
- Efficient Heteroscedastic Bayesian Optimization for Risk-Aware AutoRL, by Mingxuan Che, Tsung Yuan Tseng, Theresa Eimer, Marius Lindauer, and Alexander von Rohr.