From LLM-as-Judge to Agent-as-Judge: A Technical Evolution
The article analyzes why using a large language model alone as a judge is unreliable for complex AI evaluation, introduces the Agent-as-Judge architecture with sandboxed execution, Skill‑driven workflows, and a three‑layer assessment framework, and discusses its benefits, costs, and practical outcomes.
