Generative Reinforcement Bidding for Multi-Channel Online Ads: Tencent’s GRB Framework at KDD 2026

The article presents Tencent’s GRB (Generative Reinforcement Bidding) framework, a multi‑slot online advertising solution that combines a Decision‑Transformer‑based generative model with an Offline‑to‑On‑Policy (O2P) training paradigm and an adaptive constraint‑exploration module, and demonstrates statistically significant gains over strong baselines in both offline and large‑scale online A/B tests.

Tencent Advertising Technology
Tencent Advertising Technology
Tencent Advertising Technology
Generative Reinforcement Bidding for Multi-Channel Online Ads: Tencent’s GRB Framework at KDD 2026

Background and Motivation

Auto‑bidding is the core decision module that links advertisers’ budget and CPA targets to platform traffic allocation. Traditional academic approaches—PID control, reinforcement learning, and incremental methods—are designed for a single traffic slot, while Tencent’s real‑world scenario spans video accounts, Moments, public accounts, mini‑programs, PCAD, and alliance slots, each with distinct eCPM, conversion value, and traffic scale. This heterogeneity makes single‑slot methods infeasible.

Problem Definition

The authors formalize multi‑slot real‑time bidding as a constrained optimization problem over a fixed horizon (e.g., a day). At each decision step the agent outputs a vector of bid‑adjustment coefficients, one per slot, which are applied to all requests arriving in that step. The objective is to maximize total conversion value under budget B and target CPA constraints, while accounting for stochastic market prices and slot‑specific value distributions.

Method Design

Overall Framework : GRB consists of three key components (see Figure 3‑1):

Multi‑slot generative model – a Decision Transformer backbone with a weighted‑mask loss that jointly models all slots, addressing data sparsity and distribution imbalance.

Generative Evaluator Model (GEM) – an independently trained transformer that predicts next‑state and Return‑to‑Go (RTG) given a state‑action pair, serving as a low‑cost simulation environment.

Offline‑to‑On‑Policy (O2P) training – offline imitation learning followed by on‑policy rollouts with GEM, using a baseline‑gated replay buffer and a refinement negative log‑likelihood (RNLL) loss, thus avoiding standard RL gradients.

The multi‑slot model expands state and action representations to vectors covering all slots, concatenates slot‑specific features with a shared backbone, and feeds each slot’s representation into a lightweight expert head that outputs a bid‑adjustment distribution. Training uses a cost‑weighted NLL loss, weighting each sample by its actual spend to focus on business‑impactful slots.

Adaptive Constraint Exploration

To keep online fine‑tuning safe, the authors impose two constraints: (1) a KL‑trust‑region to limit deviation from the offline policy, and (2) a sensitivity constraint on the policy’s response to changes in the initial RTG. Two virtual queues track cumulative violations; their values dynamically adjust the constraint weights, guaranteeing that the exploration stays within a provably safe region (theoretical convergence bounds are provided).

Experiments

Main Comparison : GRB is benchmarked against five baselines (USCB, IQL, DT, GAS, GAVE) adapted to the multi‑slot setting. Across three value‑type metrics GRB consistently outperforms all baselines, with larger gains as the CPA‑penalty coefficient β increases. GRB also shows lower variance, indicating more stable performance. Ratio‑type metrics are comparable or better, confirming a superior trade‑off between value and constraint satisfaction.

Ablation Studies : Three variants—GRB‑1d (single‑slot unified bid), GRB‑E (removing the ad‑maintained replay buffer), and GRB‑A (removing the adaptive constraint module)—all perform worse than the full model, validating the necessity of joint multi‑slot modeling, the replay‑buffer strategy, and the safety module.

Evaluator Robustness : GEM’s predictions on held‑out data achieve MAPE of 6.90% (Cost) and 6.72% (GMV). Injecting ±5% noise into GEM’s output does not degrade downstream metrics, confirming the reliability of the simulation environment.

Case Study : On a representative ad, GRB’s slot‑specific adjustments increase spend on high‑value slots (e.g., public‑account +13.09% spend, CPA +5.05%) while reducing spend on low‑performing slots (e.g., alliance –52.94% spend, CPA –27.47%). Overall, total spend rises 8.75% while CPA drops 5.53%, demonstrating effective budget reallocation.

Online A/B Test : A five‑day experiment (Dec 19‑23 2025) on Tencent’s first‑day‑pay ROI game ads compares GRB against a standard unified bidding baseline. GRB achieves statistically significant improvements (p < 0.05) in spend, non‑over‑budget spend ratio, and empty‑spend ratio, indicating better budget utilization and cost control. GRB is now fully deployed in the first‑day‑pay ROI scenario and expanding to all intelligent‑placement contexts.

Conclusion

The paper tackles three challenges of multi‑slot bidding—data sparsity from coupled decisions, dominance of high‑traffic slots in training, and the offline performance ceiling—by introducing (1) a Decision‑Transformer‑based multi‑slot generative model, (2) the O2P offline‑to‑on‑policy training paradigm, and (3) an adaptive constraint‑exploration mechanism with theoretical safety guarantees. Open problems include dynamic slot topology, improving GEM’s out‑of‑distribution reliability, and enriching constraint signals beyond KL and Lipschitz.

Tencent Marketing banner
Tencent Marketing banner
Figure 1‑1: Multi‑slot eCPM‑CTR/CVR distribution
Figure 1‑1: Multi‑slot eCPM‑CTR/CVR distribution
Figure 3‑1: Multi‑slot generative method design
Figure 3‑1: Multi‑slot generative method design
Table 4‑1: Main experiment results comparing GRB with baselines
Table 4‑1: Main experiment results comparing GRB with baselines
Table 4‑2: Spend and CPA ratio comparison between GRB‑1d and GRB
Table 4‑2: Spend and CPA ratio comparison between GRB‑1d and GRB
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Decision Transformeradaptive constraint explorationgenerative reinforcement biddingmulti-slot advertisingoffline-to-on-policy
Tencent Advertising Technology
Written by

Tencent Advertising Technology

Official hub of Tencent Advertising Technology, sharing the team's latest cutting-edge achievements and advertising technology applications.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.