V-STAR: A Value‑Driven Reinforcement Learning Paradigm for Generative Recommendation
The paper identifies a structural mismatch between probability‑driven beam search and reward‑driven RL fine‑tuning in generative recommendation, proposes V-STAR with value‑guided efficient decoding (VED) and sibling‑wise GRPO to align decoding and optimization, and demonstrates superior offline and online performance through extensive experiments and ablations.
