Bandit Simulation for Average Reward Inference

Published in arXiv preprint, 2026

Multi-arm bandit algorithms are increasingly used in online platforms, clinical trials, and social science experiments, but valid statistical inference on their performance remains an open challenge. We propose Bandit Simulation for Inference (BSI), a framework that fits a simulator of the bandit environment from observed data and uses it to estimate the mean reward under evaluation policies, including adaptive black-box algorithms. BSI propagates uncertainty in the estimated simulator parameters into confidence interval construction, requires only weak exploration assumptions on the behavior policy, and avoids importance weighting.

Recommended citation: C.-Y. Chang\*, S. Praharaj\*, K. Khamaru and K. W. Zhang. "Bandit Simulation for Average Reward Inference." arXiv preprint, 2026.
Download Paper