Tagged articles

continuous evaluation

5 articles · Page 1 of 1
Woodpecker Software Testing
Woodpecker Software Testing
Sep 2, 2026 · Artificial Intelligence

Extending AI Native Application Quality Across the Full Lifecycle

This module explains how to transform AI quality assurance from a single pre‑deployment test into a continuous, organization‑wide lifecycle system, covering design‑time quality gates, CI‑integrated testing, production monitoring, safety, reliability, explainability, compliance, and a real‑world automotive case study.

AIcompliancecontinuous evaluation
0 likes · 10 min read
Extending AI Native Application Quality Across the Full Lifecycle
Alibaba Cloud Native
Alibaba Cloud Native
Sep 1, 2026 · Artificial Intelligence

From Golden Metrics to Rubric: Building a Quantifiable, Explainable Evaluation Loop for AI Agents

This article walks through constructing a fully quantifiable and explainable evaluation system for AI agents—starting with business‑level golden metrics, using LLMs to break them into a detailed Rubric, embedding the Rubric in a custom evaluator, configuring evaluation tasks with trace data, and closing the loop by turning low‑scoring cases into actionable insights for continuous improvement.

AI evaluationAgentLoopData Flywheel
0 likes · 13 min read
From Golden Metrics to Rubric: Building a Quantifiable, Explainable Evaluation Loop for AI Agents
Woodpecker Software Testing
Woodpecker Software Testing
Aug 31, 2026 · Artificial Intelligence

Shift‑Left AI Evaluation: Full Process, Practical Case Study and Best Practices

This article presents a comprehensive, six‑stage AI evaluation workflow that embeds quality checks early in development, explains the shift‑left testing philosophy, details each phase with concrete actions and metrics, and illustrates the approach with a real‑world intelligent‑customer‑service project.

AI evaluationCI/CDcontinuous evaluation
0 likes · 18 min read
Shift‑Left AI Evaluation: Full Process, Practical Case Study and Best Practices
DataFunTalk
DataFunTalk
Jul 25, 2026 · Artificial Intelligence

When Errors Spread Among AI Agents, Who Pulls the Brakes? Safe Collaborative Growth

The article analyses how self‑evolving AI agents shift from simple tool use to autonomous planning, proposes a fast‑slow thinking architecture, curriculum learning, organizational structures, continuous evaluation, and a four‑layer safety framework to ensure they grow responsibly while collaborating with humans.

AI agentsAgent SafetySelf-Evolution
0 likes · 17 min read
When Errors Spread Among AI Agents, Who Pulls the Brakes? Safe Collaborative Growth
Architect
Architect
May 3, 2026 · Artificial Intelligence

Why the Same Model Feels Different in Coding Agents: Model Sets the Capability Ceiling, Harness Sets the Production Floor

The article examines how a model defines an agent’s ultimate capabilities while the harness determines its production reliability, detailing continuous evaluation, context‑budgeting, tool‑error classification, multi‑model migration, and SRE‑style engineering practices needed to keep AI coding agents stable and performant.

AI agentsAgent HarnessContext Management
0 likes · 31 min read
Why the Same Model Feels Different in Coding Agents: Model Sets the Capability Ceiling, Harness Sets the Production Floor