Tagged articles

Rubric Design

2 articles · Page 1 of 1
dbaplus Community
dbaplus Community
Sep 15, 2026 · Artificial Intelligence

Meituan's Agent Evaluation System: From Scoring to Infrastructure Capability

Meituan's Turing team shares a comprehensive framework for evaluating AI agents, covering multi-layered assessment (result, process, efficiency, risk), human-machine alignment via binary rubrics, seed test sets, expert knowledge integration, and the shift toward infrastructure for long-horizon agents with task-based evaluation harnesses.

Agent EvaluationEvaluation InfrastructureHuman-Machine Alignment
0 likes · 33 min read
Meituan's Agent Evaluation System: From Scoring to Infrastructure Capability
FunTester
FunTester
May 16, 2026 · Artificial Intelligence

Anthropic’s Generator‑Critic Approach for Reliable Test‑Case Evaluation

The article explains why letting the same Agent both generate a test case and self‑review leads to hidden flaws, and how Anthropic’s Generator‑Critic architecture with physically isolated contexts and a well‑crafted rubric provides a more dependable way to assess test‑case quality and control retries.

Agent ArchitectureAnthropicGenerator‑Critic
0 likes · 7 min read
Anthropic’s Generator‑Critic Approach for Reliable Test‑Case Evaluation