Tagged articles

preference optimization

17 articles · Page 1 of 1
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 3, 2026 · Artificial Intelligence

OPD Evolution: From CoT SFT to Self‑Distillation and Preference Optimization

Since 2026, On‑Policy Distillation (OPD) has rapidly become a focal research area, evolving from offline teacher‑generated data to online student‑driven supervision, with advances such as OPD+, Direct OPD, weak‑to‑strong OPD, self‑distillation techniques, and preference‑optimization signals reshaping post‑training for large language models.

Large Language ModelsNLPOPD
0 likes · 7 min read
OPD Evolution: From CoT SFT to Self‑Distillation and Preference Optimization
DaTaobao Tech
DaTaobao Tech
Jul 3, 2026 · Artificial Intelligence

LocalDPO: A CVPR 2026 Method for Fine‑Grained Preference Optimization in Video Diffusion Models

LocalDPO introduces a zero‑annotation, region‑aware DPO framework that uses high‑quality real videos as positive samples and automatically generated locally degraded negatives to align video diffusion models with human preferences, achieving significant gains in visual quality, temporal consistency, and subjective ratings on CogVideoX and Wan2.1.

Artificial IntelligenceCVPR 2026LocalDPO
0 likes · 13 min read
LocalDPO: A CVPR 2026 Method for Fine‑Grained Preference Optimization in Video Diffusion Models
JD Retail Technology
JD Retail Technology
Jun 23, 2026 · Artificial Intelligence

How RAD‑DPO Aligns Preferences in OxygenSearch Generative Retrieval (SIGIR 2026)

This article analyzes the challenges of generative retrieval for e‑commerce (shared SID prefixes, noisy implicit feedback, and probability squeezing) and presents RAD‑DPO, a robust adaptive denoising direct preference optimization that uses session‑level multi‑label contrast, token‑level gradient detachment, and dynamic reward weighting to improve both effectiveness and training efficiency.

RAD-DPOe-commerce searchgenerative retrieval
0 likes · 18 min read
How RAD‑DPO Aligns Preferences in OxygenSearch Generative Retrieval (SIGIR 2026)
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jun 21, 2026 · Artificial Intelligence

Rank‑Only Rewards Accelerate One‑Step Text‑to‑Image Preference Optimization 3.5×

DrPO introduces a drifting‑field based, rank‑only reward mechanism for one‑step text‑to‑image models, enabling reinforcement‑learning‑after‑training without back‑propagating reward gradients; it speeds up training 3.51× versus DRaFT, works with non‑differentiable rewards, and improves generation quality on SD‑Turbo and SDXL‑Turbo.

DrPODrifting ModelHPSv3
0 likes · 11 min read
Rank‑Only Rewards Accelerate One‑Step Text‑to‑Image Preference Optimization 3.5×
Data Party THU
Data Party THU
May 30, 2026 · Artificial Intelligence

How USTC’s Tiny LCPO Training Cuts Large Model Overthinking in Half

The paper introduces LCPO, a lightweight preference‑optimization technique that uses only 800 training examples and 50 steps to teach large language models to produce concise, accurate answers, halving inference length while often improving accuracy and reducing training cost by up to two orders of magnitude.

LCPOLarge Language ModelsLow-Resource Training
0 likes · 8 min read
How USTC’s Tiny LCPO Training Cuts Large Model Overthinking in Half
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
May 20, 2026 · Artificial Intelligence

How 800 Data Points Halve LLM Chain‑of‑Thought Length and Boost Accuracy

The ICLR‑2026 paper introduces LCPO, a lightweight preference‑optimization technique that uses only 800 curated examples and 50 training steps to cut large‑model chain‑of‑thought generation length by about 50% while maintaining or even improving answer accuracy, dramatically reducing training and inference costs.

LCPOLarge Language ModelsLow-Resource Training
0 likes · 8 min read
How 800 Data Points Halve LLM Chain‑of‑Thought Length and Boost Accuracy
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
May 14, 2026 · Artificial Intelligence

Turning Multi‑Teacher Conflict into Dynamic Constraints: Robust Reasoning Alignment for Multimodal LLMs (ICML 2026)

APO (Autonomous Preference Optimization) converts the drift and conflict among multiple teacher multimodal LLMs into dynamic negative constraints while treating consensus as a positive preference, enabling robust concept alignment and superior diagnostic accuracy on the CXR‑MAX benchmark, as demonstrated by extensive ICML‑2026 experiments.

APOICML 2026Multimodal LLM
0 likes · 11 min read
Turning Multi‑Teacher Conflict into Dynamic Constraints: Robust Reasoning Alignment for Multimodal LLMs (ICML 2026)
AI Frontier Lectures
AI Frontier Lectures
Jan 21, 2026 · Artificial Intelligence

How AP2O‑Coder Cuts LLM Code Errors by Up to 3% with Adaptive Preference Optimization

The paper introduces AP2O‑Coder, an adaptive progressive preference optimization framework that systematically captures error types, progressively refines LLM code generation, and dynamically adapts training data, achieving up to a 3% pass@k improvement across multiple open‑source models while reducing data requirements.

AP2O-CoderLLMSoftware Engineering
0 likes · 11 min read
How AP2O‑Coder Cuts LLM Code Errors by Up to 3% with Adaptive Preference Optimization
Kuaishou Tech
Kuaishou Tech
Dec 3, 2025 · Artificial Intelligence

Can Diffusion Models Be Their Own Reward Model? Latent Reward Modeling & Step-Level Preference Optimization

This article presents a novel paradigm—Latent Reward Model (LRM) and Latent Preference Optimization (LPO)—that repurposes diffusion models as noise‑aware latent reward models for step‑level preference optimization, addressing the shortcomings of pixel‑level reward models, introducing multi‑preference consistent filtering, and demonstrating significant performance and efficiency gains on benchmarks such as PickScore and T2I‑CompBench++.

AI Alignmentdiffusion modelsimage generation
0 likes · 9 min read
Can Diffusion Models Be Their Own Reward Model? Latent Reward Modeling & Step-Level Preference Optimization
Meituan Technology Team
Meituan Technology Team
Jul 31, 2025 · Artificial Intelligence

8 Must-Read ACL 2025 Papers from Meituan: Generative Retrieval, Multimodal LLMs & More

Meituan’s research team showcases eight ACL 2025 papers spanning generative retrieval, multi‑objective preference alignment, rich‑text image understanding, cross‑language transfer, multimodal math reasoning, and more, offering insights and breakthroughs that can inspire and aid fellow researchers.

ACL 2025Code-SwitchingMultimodal LLM
0 likes · 15 min read
8 Must-Read ACL 2025 Papers from Meituan: Generative Retrieval, Multimodal LLMs & More
JD Tech Talk
JD Tech Talk
Mar 13, 2025 · Artificial Intelligence

CTR-Driven Advertising Image Generation with Multimodal Large Language Models

This paper proposes CAIG, a novel method for generating high-CTR advertising images using multimodal large language models, combining reinforcement learning and preference optimization to align generated content with product features.

CTR predictionadvertising image generationmultimodal large language models
0 likes · 10 min read
CTR-Driven Advertising Image Generation with Multimodal Large Language Models
DataFunSummit
DataFunSummit
Nov 28, 2024 · Artificial Intelligence

Generative Retrieval for E‑commerce Search: Lexical and SemanticID Approaches

This article presents a comprehensive study of generative retrieval for large‑scale e‑commerce search, detailing background challenges, the advantages of generative methods, two concrete strategies—Lexical‑based and SemanticID‑based—along with task redesign, preference optimization, constrained beam search, extensive experiments, and future research directions.

e-commerce searchgenerative retrievallexical approach
0 likes · 21 min read
Generative Retrieval for E‑commerce Search: Lexical and SemanticID Approaches
Bilibili Tech
Bilibili Tech
Nov 5, 2024 · Artificial Intelligence

Bilibili's In-House Role-Playing Large Language Model: Architecture, Training Stages, Evaluation, and Demonstrations

Bilibili’s in‑house role‑playing large language model, built on the Index architecture and refined through pre‑training, supervised fine‑tuning, and preference optimization (PPO and DPO), achieved top scores on the Chinese CharacterEval benchmark, surpassing rivals while incorporating safety alignment and showcasing consistent, personality‑driven dialogue examples.

Supervised Fine‑Tuningcontent safetyevaluation benchmark
0 likes · 13 min read
Bilibili's In-House Role-Playing Large Language Model: Architecture, Training Stages, Evaluation, and Demonstrations
NewBeeNLP
NewBeeNLP
Aug 7, 2024 · Artificial Intelligence

Can Intuitive Fine‑Tuning Replace Expensive RLHF and DPO for LLM Alignment?

This article analyses the shortcomings of current large language model training methods such as SFT, RLHF and DPO, explains why they incur high data and compute costs, and introduces Intuitive Fine‑Tuning (IFT) with temporal residual connections as a cheaper yet effective alternative that better aligns training objectives with real generation tasks.

DPOIntuitive Fine-TuningLLM
0 likes · 15 min read
Can Intuitive Fine‑Tuning Replace Expensive RLHF and DPO for LLM Alignment?
NewBeeNLP
NewBeeNLP
May 13, 2024 · Artificial Intelligence

Why DPO Treats LLMs as Q‑Functions: A Deep Theoretical Dive

This article offers a detailed theoretical interpretation of the DPO algorithm, showing how large language models can be viewed as Q‑functions, unifying sequence‑wise and step‑wise decision perspectives, and discussing the resulting implications for reinforcement‑learning‑based alignment research.

DPOLLMQ-Function
0 likes · 14 min read
Why DPO Treats LLMs as Q‑Functions: A Deep Theoretical Dive