Tagged articles

AI Agent evaluation

5 articles · Page 1 of 1
ThinkingAgent
ThinkingAgent
Oct 8, 2026 · Artificial Intelligence

No Evals, No Production Agents: The 4-Layer Evaluation Framework

This article presents a comprehensive four-layer evaluation framework for production-grade AI agents—Model, Component, Trajectory, and Outcome Eval—along with three dataset types (Golden, Edge Case, Adversarial), three evaluation methods (Rule-based, LLM-as-Judge, Human), seven key metrics, regression automation, and a production-to-eval flywheel, illustrated with a customer-service refund agent case study.

AI Agent evaluationAgent ObservabilityAgent Security
0 likes · 36 min read
No Evals, No Production Agents: The 4-Layer Evaluation Framework
FunTester
FunTester
Sep 27, 2026 · Artificial Intelligence

Why AI Agents Silently Fail: A Three-Layer Evaluation Framework

This article explains why AI Agents produce plausible but incorrect outputs, introduces a three-layer evaluation framework (deterministic rules, fact verification, model-based quality review), and provides a practical one-week startup plan using golden datasets, isolated test runs, and score thresholds to ensure reliable Agent outputs.

AI Agent evaluationAI reliabilitydeterministic rules
0 likes · 14 min read
Why AI Agents Silently Fail: A Three-Layer Evaluation Framework
Software Engineering 3.0 Era
Software Engineering 3.0 Era
Sep 7, 2026 · Artificial Intelligence

AI Agent Evaluation Guide: Building Observable, Evaluable, Self-Evolving Quality Systems

This comprehensive guide synthesizes 2026 industry practices from Xiaohongshu and Alipay to build production-ready AI Agent evaluation systems, covering metrics (Quality/Cost/Safety), three-tier evaluation granularities, Judge system design, OpenTelemetry-based observability, platform architecture with contract-driven test generation, dual flywheel offline/online loops, and self-evolving prompt optimization — moving evaluation from post-hoc verification to embedded engineering guardrails.

AI Agent evaluationAgentOpsLLM-as-Judge
0 likes · 37 min read
AI Agent Evaluation Guide: Building Observable, Evaluable, Self-Evolving Quality Systems
Machine Heart
Machine Heart
Mar 31, 2026 · Artificial Intelligence

What Does DeepResearch Bench Measure? Toward Human‑Level AI Agent Evaluation

The DeepResearch Bench and Bench II, open‑source benchmarks from the USTC team, evaluate deep‑research AI agents on report quality, citation reliability, and information recall using the RACE and FACT frameworks, aiming to align automated scores with human expert judgments.

AI Agent evaluationDeepResearch BenchFACT
0 likes · 12 min read
What Does DeepResearch Bench Measure? Toward Human‑Level AI Agent Evaluation