Tagged articles

Evaluation Methodology

3 articles · Page 1 of 1
Software Engineering 3.0 Era
Software Engineering 3.0 Era
Sep 7, 2026 · Artificial Intelligence

AI Agent Evaluation Guide: Building Observable, Evaluable, Self-Evolving Quality Systems

This comprehensive guide synthesizes 2026 industry practices from Xiaohongshu and Alipay to build production-ready AI Agent evaluation systems, covering metrics (Quality/Cost/Safety), three-tier evaluation granularities, Judge system design, OpenTelemetry-based observability, platform architecture with contract-driven test generation, dual flywheel offline/online loops, and self-evolving prompt optimization — moving evaluation from post-hoc verification to embedded engineering guardrails.

AI Agent EvaluationAgentOpsEvaluation Methodology
0 likes · 37 min read
AI Agent Evaluation Guide: Building Observable, Evaluable, Self-Evolving Quality Systems
Shi's AI Notebook
Shi's AI Notebook
Apr 23, 2026 · Artificial Intelligence

Decoding Anthropic’s Agent Evaluation Methodology: Challenges, Graders, and Best Practices

Anthropic’s engineering blog outlines a systematic approach to evaluating AI agents, highlighting why agents are harder to test than traditional software, defining key concepts like tasks, trials, transcripts, and outcomes, and detailing the three grader types, evaluation timing, and practical decisions for building robust eval pipelines.

AI agentsEvaluation MethodologyLLM-as-Judge
0 likes · 23 min read
Decoding Anthropic’s Agent Evaluation Methodology: Challenges, Graders, and Best Practices
AI Info Trend
AI Info Trend
Oct 28, 2025 · Industry Insights

2025: The AI Agent Year and a New Standard to End the Evaluation Black Box

In 2025, China’s AI strategy targets over 90% adoption of AI agents, yet enterprises struggle with selection, acceptance, and optimization due to a lack of unified performance metrics, prompting the first national group standard—‘Enterprise‑level AI Agent Application Performance Evaluation Specification’—to provide a comprehensive, multi‑dimensional assessment framework for developers, users, and third‑party evaluators.

AI agentsEvaluation Methodologyartificial intelligence
0 likes · 7 min read
2025: The AI Agent Year and a New Standard to End the Evaluation Black Box