This is the pre-proceedings for the RLC 2026. You may expect minor changes.

Bandits for Efficient Experimentation: Adapting to Control Group, Preferences, and Context Drifts

By Udvas Das, Waris Radji, Debabrota Basu, and Odalric-Ambrym Maillard

Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

Will be presented at the Reinforcement Learning Conference (RLC), MontrĂ©al, Quebec, Canada, August 15–17, 2026.


Download:

Keywords: Heteroskedastic Bandits, Context-drift, Customised preference, Control group,

Abstract:

We consider a variant of the linear contextual stochastic multi-armed bandits, where the learner must provide recommendations to a group of users, each having its personalized preference vector, and in the presence of context distributions that are drifting over time. Under practitioner-friendly assumptions, we reduce this setting to linear bandit with stationary mean but heteroscedastic and non-stationary noise. We further study the case when the learner must ensure the mean reward of each decision must exceed that of a baseline strategy $\boldsymbol{\pi}_0$ at each decision step. We introduce **Dri-MED**, an algorithm inspired from the linear version of the MED strategy, and carefully adapted to handle the non-stationary heteroskedastic noise. We show that the instance-dependent regret scales as $\tilde{\mathcal O}\left(\frac{\kappa}{\tilde{\Delta}}d^2(\log(T)\right)$, where $\tilde{\Delta}$ is the constraint-aware sub-optimality gap subject to policy $\pi_0$, with variance-aware multiplicative term $\kappa$ that we carefully handle using heteroscedastic regression. We further show **Dri-MED** enjoys $\tilde{\mathcal O}(d)$ expected constraint violations. Our numerical results suggest that \texttt{Dri-MED} significantly outperforms conservative baselines that ignores the drift and preference structure.


Citation Information:

Udvas Das, Waris Radji, Debabrota Basu, and Odalric-Ambrym Maillard. "Bandits for Efficient Experimentation: Adapting to Control Group, Preferences, and Context Drifts." Reinforcement Learning Journal, vol. 7, 2026, pp. TBD.

BibTeX:
@article{das2026bandits,
    title={Bandits for Efficient Experimentation: Adapting to Control Group, Preferences, and Context Drifts},
    author={Udvas Das and Waris Radji and Debabrota Basu and Odalric-Ambrym Maillard},
    journal={Reinforcement Learning Journal},
    volume={7},
    pages={},
    year={2026}
}