Tagged articles

preference alignment

5 articles · Page 1 of 1
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Sep 12, 2026 · Artificial Intelligence

Robots Retracing LLMs' Scaling Path: LightNav-0 & Light REACT Explained

Light Source Innovation, founded by ex-OpenAI RLHF expert Jiang Xu, releases LightNav-0 for zero-shot cross-morphology navigation and Light REACT for whole-body resilience control, applying LLM-style scalable pre-training, alignment, and deployment paradigms to embodied AI with sim-to-real synthetic data and preference-aligned RL.

Embodied AIRLHFpreference alignment
0 likes · 14 min read
Robots Retracing LLMs' Scaling Path: LightNav-0 & Light REACT Explained
Data Party THU
Data Party THU
May 23, 2026 · Artificial Intelligence

ProteinOPD: Tsinghua’s Efficient Multi‑Objective Preference Alignment Framework for Protein Design

ProteinOPD introduces a multi‑teacher, on‑policy preference‑distillation framework that aligns protein language models with multiple design objectives—foldability, solubility and thermostability—while preserving generation quality, achieving up to 54% stability gains and an eight‑fold training speedup.

ProteinOPDdeep learninglanguage models
0 likes · 9 min read
ProteinOPD: Tsinghua’s Efficient Multi‑Objective Preference Alignment Framework for Protein Design
Weekly Large Model Application
Weekly Large Model Application
May 5, 2026 · Artificial Intelligence

What Do End‑to‑End Speech Large Models Actually Learn? A Four‑Step Diagram

The article distinguishes two meanings of “end‑to‑end,” then outlines four sequential stages—defining data and scenario, massive pre‑training on audio‑text pairs, task alignment via instruction or supervised fine‑tuning, and optional preference tuning—to guide engineers in building usable speech assistants.

PretrainingSpeech AIaudio data
0 likes · 6 min read
What Do End‑to‑End Speech Large Models Actually Learn? A Four‑Step Diagram
Weekly Large Model Application
Weekly Large Model Application
May 5, 2026 · Artificial Intelligence

Understanding Preference Alignment: Why Voice Output Needs an Extra Layer

The article explains that after task alignment, teams can produce functional demos, but true competitiveness requires preference alignment—optimizing for human comfort across dimensions like brevity, tone, and safety—and discusses how RLHF and DPO address this, especially the additional challenges of generating natural, responsive voice output.

AI AlignmentDPOHuman Feedback
0 likes · 7 min read
Understanding Preference Alignment: Why Voice Output Needs an Extra Layer