Online and Off-Policy Learning for Large Action Spaces: Structured Exploration & Optimization
This PhD thesis addresses large action space contextual bandits, proposing mixed-effects and diffusion Thompson sampling for online learning, and structured direct methods, policy-weighted likelihood, exponential smoothing, and PAC-Bayes pessimism for offline learning, showing that action structure, optimizable objectives, and pessimism are crucial for scalable decision-making.
