How to Master Online and Offline Policy Learning in Massive Action Spaces
This article reviews a PhD thesis that systematically studies online and offline learning for contextual bandits with huge action spaces, highlighting statistical, computational, and optimization challenges and presenting mixed‑effect Thompson sampling, diffusion priors, structured direct methods, and PAC‑Bayes pessimism as effective solutions.
