Tencent Advertising Technology
Aug 8, 2026 · Artificial Intelligence
V-STAR: A Value‑Driven Reinforcement Learning Paradigm for Generative Recommendation
The paper identifies a structural mismatch between probability‑driven beam search and reward‑driven RL fine‑tuning in generative recommendation, proposes V-STAR with value‑guided efficient decoding (VED) and sibling‑wise GRPO to align decoding and optimization, and demonstrates superior offline and online performance through extensive experiments and ablations.
Reinforcement LearningV‑STARgenerative recommendation
0 likes · 13 min read
