Redefining Agent Evaluation: An Offline LLM-as-a-Judge Framework and Practice
The article presents a quantitative offline evaluation system for AI agents in real outbound-call scenarios, combining reference‑based scoring with pairwise GSB methods, addressing regression and optimization, mitigating systematic bias through judge model selection and majority‑vote adjustments, and delivering an objective, high‑efficiency benchmark.
