Tagged articles

Rubrics

5 articles · Page 1 of 1
AntData
AntData
Jul 31, 2026 · Artificial Intelligence

When Expert Experience Can Be Quantified: How Rubrics Become Data Assets for LLM Inference Training

The article analyzes how combining formal verification with expert‑derived Rubrics provides fine‑grained process supervision for large language models, presents the CRAFT data‑production pipeline, and shows experimental gains on math and medical benchmarks using Rubric‑driven RL, SFT, and alternating RL‑SFT training.

Formal VerificationLLMRubrics
0 likes · 23 min read
When Expert Experience Can Be Quantified: How Rubrics Become Data Assets for LLM Inference Training
Data Party THU
Data Party THU
Jun 27, 2026 · Artificial Intelligence

Defining a Good Answer in the Agent Era: A Rubrics Survey

This survey examines how rubrics—structured, multi‑dimensional evaluation criteria—are defined, constructed, and applied to train and evaluate large language models, especially for open‑ended, high‑risk and agentic tasks, while highlighting current challenges such as reward hacking and bias.

AI safetyAgentLarge Language Models
0 likes · 15 min read
Defining a Good Answer in the Agent Era: A Rubrics Survey
DataFunSummit
DataFunSummit
Jun 9, 2026 · Artificial Intelligence

From Gut Feelings to Measurable Metrics: Practicing the Rubrics‑Based Expert Knowledge Extraction and Annotation System CRAFT

The article analyzes the growing difficulty of evaluating large AI models, critiques traditional RLVR and RLHF approaches, introduces a Rubrics‑based evaluation paradigm, describes the design and three‑stage workflow of the CRAFT system, reports math‑domain experiments showing up to 6.2 percentage‑point gains, and outlines future extensions to other domains.

AI evaluationCRAFTLarge Language Models
0 likes · 14 min read
From Gut Feelings to Measurable Metrics: Practicing the Rubrics‑Based Expert Knowledge Extraction and Annotation System CRAFT
PaperAgent
PaperAgent
Jun 9, 2026 · Artificial Intelligence

Defining Standard Answers for Agent‑Era LLMs: A Rubrics Survey

The survey from RUC‑Gaoling AI Institute reviews Rubrics for large language models, explaining why they are needed for open‑ended, high‑risk tasks, how they are constructed, and how they can be applied to policy and reward model training as well as multi‑dimensional evaluation across general and domain‑specific scenarios.

AgentLLMRubrics
0 likes · 14 min read
Defining Standard Answers for Agent‑Era LLMs: A Rubrics Survey
Machine Heart
Machine Heart
May 31, 2026 · Artificial Intelligence

Defining a Good Answer in the Agent Era: A Rubrics Survey

This survey examines how rubrics can decompose the vague notion of a "good answer" for large language models into concrete, multi‑dimensional evaluation criteria, detailing their definition, construction methods, applications in training and evaluation, and the open challenges they present.

AI alignmentAgentic AILarge Language Models
0 likes · 13 min read
Defining a Good Answer in the Agent Era: A Rubrics Survey