Tagged articles

Rubric

6 articles · Page 1 of 1
Machine Heart
Machine Heart
Sep 23, 2026 · Industry Insights

Muchen AI: Redefining AI Data Infrastructure Through Verifiable Evaluation Standards

Muchen AI moves beyond data labeling to build long-horizon evaluation systems using structured Rubrics, Docker-based reproducible environments, and automated scoring, raising evaluator consistency from 30% to 90% for code models and extending the framework to scientific research via ScienceBuddy's recursive verification loop across 10+ domains.

AI data engineeringAI infrastructureRubric
0 likes · 13 min read
Muchen AI: Redefining AI Data Infrastructure Through Verifiable Evaluation Standards
Alibaba Cloud Native
Alibaba Cloud Native
Sep 1, 2026 · Artificial Intelligence

From Golden Metrics to Rubric: Building a Quantifiable, Explainable Evaluation Loop for AI Agents

This article walks through constructing a fully quantifiable and explainable evaluation system for AI agents—starting with business‑level golden metrics, using LLMs to break them into a detailed Rubric, embedding the Rubric in a custom evaluator, configuring evaluation tasks with trace data, and closing the loop by turning low‑scoring cases into actionable insights for continuous improvement.

AI evaluationAgentLoopData Flywheel
0 likes · 13 min read
From Golden Metrics to Rubric: Building a Quantifiable, Explainable Evaluation Loop for AI Agents
Woodpecker Software Testing
Woodpecker Software Testing
Aug 23, 2026 · Artificial Intelligence

How to Choose the Right Model Evaluation Method – From Exact Match to LLM-as-a-Judge

This guide explains objective metrics such as Exact and Fuzzy Match, the QUEST framework for human evaluation, rubric design and calibration, the LLM-as-a-Judge approach with its biases and trade‑offs, and a five‑dimensional evaluation framework for building robust, explainable and fair AI systems.

AI assessmentLLM-as-a-JudgeRubric
0 likes · 26 min read
How to Choose the Right Model Evaluation Method – From Exact Match to LLM-as-a-Judge
Meituan Technology Team
Meituan Technology Team
Aug 6, 2026 · Artificial Intelligence

A Deep Dive into Agent Evaluation: From Basics to Advanced Practices

This article explains why evaluating AI agents requires more than answer correctness, outlines a four‑layer evaluation framework (result, process, efficiency, risk), compares short‑ and long‑horizon agents, and presents a practical methodology that combines objective and subjective metrics, rubric binary‑ization, case management, and infrastructure requirements for scalable, repeatable agent testing.

AI AgentAgent EvaluationRubric
0 likes · 25 min read
A Deep Dive into Agent Evaluation: From Basics to Advanced Practices
FunTester
FunTester
May 17, 2026 · Artificial Intelligence

How a Rubric‑Driven Agent Achieves More Stable Outputs

The article explains why vague expectations cause unstable Agent results, introduces Rubric as a concrete, pre‑written scoring standard for Generator‑Critic workflows, details how to design clear Yes/No criteria, organize them into Must/Should/Nice‑to‑have layers, and iteratively refine the Rubric for reliable AI output.

AI evaluationAgentCritic
0 likes · 8 min read
How a Rubric‑Driven Agent Achieves More Stable Outputs