Tagged articles

Huber Loss

1 articles · Page 1 of 1
PaperAgent
PaperAgent
Jul 9, 2026 · Artificial Intelligence

A New Paradigm for LLM Reward Modeling: Mixing Huber and Hinge Losses in E‑GRM

The article analyzes the E‑GRM framework's need for both accurate score regression and stable ranking signals, proposes a weighted combination of Huber and hinge losses, and demonstrates through extensive ablations and downstream GRPO experiments that the mixed loss yields superior calibration, ranking, and policy‑learning performance.

E‑GRMHinge LossHuber Loss
0 likes · 10 min read
A New Paradigm for LLM Reward Modeling: Mixing Huber and Hinge Losses in E‑GRM