JD Retail Technology
Jun 23, 2026 · Artificial Intelligence
How RAD‑DPO Aligns Preferences in OxygenSearch Generative Retrieval (SIGIR 2026)
This article analyzes the challenges of generative retrieval for e‑commerce (shared SID prefixes, noisy implicit feedback, and probability squeezing) and presents RAD‑DPO, a robust adaptive denoising direct preference optimization that uses session‑level multi‑label contrast, token‑level gradient detachment, and dynamic reward weighting to improve both effectiveness and training efficiency.
Preference OptimizationRAD-DPOe-commerce search
0 likes · 18 min read
