Tagged articles

E‑GRM

2 articles · Page 1 of 1
PaperAgent
PaperAgent
Jul 9, 2026 · Artificial Intelligence

A New Paradigm for LLM Reward Modeling: Mixing Huber and Hinge Losses in E‑GRM

The article analyzes the E‑GRM framework's need for both accurate score regression and stable ranking signals, proposes a weighted combination of Huber and hinge losses, and demonstrates through extensive ablations and downstream GRPO experiments that the mixed loss yields superior calibration, ranking, and policy‑learning performance.

E‑GRMHinge LossHuber Loss
0 likes · 10 min read
A New Paradigm for LLM Reward Modeling: Mixing Huber and Hinge Losses in E‑GRM
Data Party THU
Data Party THU
Jul 7, 2026 · Artificial Intelligence

Parallel Decoding for Large Language Models: Balancing Inference Speed and Sampling Diversity in E‑GRM

The article presents an engineering analysis of the E‑GRM framework, detailing how parallel decoding, a temperature‑ladder sampling strategy, and batch‑parallel KV‑Cache sharing achieve low‑latency, high‑diversity inference while preserving consensus‑driven routing accuracy.

Batch ParallelismConsensus RoutingE‑GRM
0 likes · 14 min read
Parallel Decoding for Large Language Models: Balancing Inference Speed and Sampling Diversity in E‑GRM