Tagged articles

Thompson sampling

8 articles · Page 1 of 1
Data Party THU
Data Party THU
Sep 9, 2026 · Artificial Intelligence

Online and Off-Policy Learning for Large Action Spaces: Structured Exploration & Optimization

This PhD thesis addresses large action space contextual bandits, proposing mixed-effects and diffusion Thompson sampling for online learning, and structured direct methods, policy-weighted likelihood, exponential smoothing, and PAC-Bayes pessimism for offline learning, showing that action structure, optimizable objectives, and pessimism are crucial for scalable decision-making.

PAC-BayesThompson samplingcontextual bandits
0 likes · 21 min read
Online and Off-Policy Learning for Large Action Spaces: Structured Exploration & Optimization
DeepHub IMBA
DeepHub IMBA
Sep 3, 2026 · Artificial Intelligence

FlashSpec: Adaptive LLM Inference with Speculative Decoding — Six Hard-Won Lessons

FlashSpec implements speculative decoding using GPU-native Triton kernel verification and online bandit-based draft model selection, sharing six practical lessons on specification-first development, hidden temperature bugs, cross-platform packaging pitfalls, kernel performance trade-offs, property-based testing value, and adaptive algorithm prerequisites.

CI/CDLLM inferenceThompson sampling
0 likes · 15 min read
FlashSpec: Adaptive LLM Inference with Speculative Decoding — Six Hard-Won Lessons
Data Party THU
Data Party THU
Aug 18, 2026 · Artificial Intelligence

How to Master Online and Offline Policy Learning in Massive Action Spaces

This article reviews a PhD thesis that systematically studies online and offline learning for contextual bandits with huge action spaces, highlighting statistical, computational, and optimization challenges and presenting mixed‑effect Thompson sampling, diffusion priors, structured direct methods, and PAC‑Bayes pessimism as effective solutions.

PAC-BayesThompson samplingcontextual bandits
0 likes · 18 min read
How to Master Online and Offline Policy Learning in Massive Action Spaces
JD Retail Technology
JD Retail Technology
Aug 10, 2026 · Artificial Intelligence

Bayesian Ensemble (BE): Adaptive Ensemble Selection for Bandits and Reinforcement Learning

The paper introduces Bayesian Ensemble (BE), a lightweight Bayesian layer that dynamically updates the sampling distribution over ensemble members using observed rewards, and extends it to Bayesian Ensemble Bandit (BEB) and Bayesian Ensemble DQN (BE‑DQN), achieving significant regret and click‑through improvements across synthetic, real‑world, and RL benchmarks with minimal computational overhead.

BanditBayesian EnsembleEnsemble Methods
0 likes · 12 min read
Bayesian Ensemble (BE): Adaptive Ensemble Selection for Bandits and Reinforcement Learning
DataFunSummit
DataFunSummit
Jul 30, 2026 · Industry Insights

What an 8.8 M‑User Experiment Shows About the True Limits of Agent‑Based Marketing

A longitudinal study of 8.8 million users over 11 months demonstrates that while agentic marketing can sustain performance after human optimization stops, its gains eventually decay, highlighting the emerging “Agentic CDP” model where AI handles baseline decisions while humans drive growth.

AI PersonalizationAgentic MarketingCustomer Data Platform
0 likes · 15 min read
What an 8.8 M‑User Experiment Shows About the True Limits of Agent‑Based Marketing
HomeTech
HomeTech
Jun 10, 2020 · Artificial Intelligence

Exploitation & Exploration Algorithms in Recommender Systems: ε‑Greedy, UCB, and Thompson Sampling Applications

This article introduces recommender systems and the exploitation‑exploration dilemma, explains common E&E algorithms such as ε‑greedy, Upper‑Confidence‑Bound, and Thompson Sampling, and details their practical deployment for interest‑point eviction, selection, and adaptive recall count optimization in an automotive recommendation platform.

Bandit AlgorithmsEpsilon-GreedyExploitation
0 likes · 10 min read
Exploitation & Exploration Algorithms in Recommender Systems: ε‑Greedy, UCB, and Thompson Sampling Applications
DataFunTalk
DataFunTalk
Apr 19, 2020 · Artificial Intelligence

Bandit Algorithms for Recommendation Systems: Context‑Free, Thompson Sampling, and Contextual Approaches

This article explains how multi‑armed bandit methods such as Upper Confidence Bound, Thompson Sampling, and their contextual extensions can address cold‑start, diversity, and bias problems in large‑scale recommendation systems, describing practical update mechanisms, offline evaluation techniques, and deployment experiences at Ctrip.

AIBandit AlgorithmsExploration‑exploitation
0 likes · 15 min read
Bandit Algorithms for Recommendation Systems: Context‑Free, Thompson Sampling, and Contextual Approaches
21CTO
21CTO
Jul 1, 2017 · Product Management

Why Simple Click Counts Fail: Smarter Scoring Strategies for Content Recommendation

The article recounts a junior engineer's journey improving a news app's recommendation system, moving from naive click counts to recent clicks, CTR, lower confidence bounds, and advanced multi‑armed bandit techniques like UCB and Thompson Sampling to balance relevance and novelty.

CTRLCBThompson sampling
0 likes · 9 min read
Why Simple Click Counts Fail: Smarter Scoring Strategies for Content Recommendation