Tagged articles

Skill‑RM

1 articles · Page 1 of 1
PaperAgent
PaperAgent
Jun 11, 2026 · Artificial Intelligence

Skill‑RM Shows More Resources Can Harm LLM Scoring – A Deep Dive into Alibaba’s New Evaluation Framework

The Skill‑RM paper reveals that simply appending evaluation resources can degrade large‑model scoring, while structuring those resources into a Reward‑Evaluation Skill boosts performance across benchmarks, best‑of‑N selection, and RL‑based instruction following.

Alibaba QwenEvaluation FrameworkLarge Language Models
0 likes · 7 min read
Skill‑RM Shows More Resources Can Harm LLM Scoring – A Deep Dive into Alibaba’s New Evaluation Framework