Tagged articles

Agent Evaluation

25 articles · Page 1 of 1
DataFunSummit
DataFunSummit
Sep 28, 2026 · Artificial Intelligence

10 Financial Firms Share AI Agent Strategies for 98.5% Auto-Review, 0.003% Fraud

This article analyzes how 10 leading financial institutions implement AI agents in low-tolerance scenarios, detailing their approaches to data ontology, semantic layers, multi-agent architectures, risk control, and evaluation frameworks, achieving metrics like 98.5% automated review rates and 0.003% fraud rates while ensuring auditability and regulatory compliance.

AI agentsAgent EvaluationApache Ossie
0 likes · 43 min read
10 Financial Firms Share AI Agent Strategies for 98.5% Auto-Review, 0.003% Fraud
Machine Heart
Machine Heart
Sep 26, 2026 · Artificial Intelligence

Right Answer, Wrong Reason: LexAgentHallu Benchmarks Hidden Hallucinations in Legal AI Agents

HKUST researchers introduce LexAgentHallu, a hierarchical benchmark that evaluates hallucinations in legal AI agents by tracing entire reasoning trajectories, revealing that even top-performing systems exhibit hallucinations in 89% of execution traces and 68% of correct answers contain flawed reasoning.

Agent EvaluationHKUSTHallucination Detection
0 likes · 14 min read
Right Answer, Wrong Reason: LexAgentHallu Benchmarks Hidden Hallucinations in Legal AI Agents
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Sep 21, 2026 · Artificial Intelligence

LongDS v1.1 Benchmark Released: GPT-6 Astra Tops Long-Horizon Data Analysis Lite Leaderboard

Zhejiang University and Ant Group release LongDS v1.1, a benchmark for long-horizon multi-turn data analysis agents built from real Kaggle workflows; the Lite subset of 24 tasks and 777 turns shows GPT-6 Astra leading at 78.17, with detailed error analysis revealing cascading state errors as the primary failure mode.

Agent EvaluationEMNLP 2026GPT-6 Astra
0 likes · 24 min read
LongDS v1.1 Benchmark Released: GPT-6 Astra Tops Long-Horizon Data Analysis Lite Leaderboard
Su San Talks Tech
Su San Talks Tech
Sep 21, 2026 · Artificial Intelligence

AI Agent Interview Deep Dive: 10 Critical Questions from Architecture to Evaluation

This comprehensive guide covers 10 essential AI Agent interview topics, including Agent vs LLM differences, Workflow vs Agent selection, reasoning paradigms, Function Calling, MCP, error handling, memory management, context optimization, RAG pipelines, and evaluation metrics, with code examples and architectural diagrams.

AI AgentAgent EvaluationContext Management
0 likes · 43 min read
AI Agent Interview Deep Dive: 10 Critical Questions from Architecture to Evaluation
DataFunSummit
DataFunSummit
Sep 18, 2026 · Artificial Intelligence

Palantir's AI FDE Automates Forward Deployment While China Still Recruits Human FDEs

Palantir's AI Forward Deployed Engineer (FDE) now automates execution tasks like data integration and ontology management within Foundry, while AIP Evolve enables agents to self-optimize models and prompts via eval-driven loops; Chinese enterprises similarly adopt AI FDE to parallelize delivery workflows, but human judgment remains essential for defining context, correctness, and error boundaries.

AI FDEAIP EvalsAIP Evolve
0 likes · 21 min read
Palantir's AI FDE Automates Forward Deployment While China Still Recruits Human FDEs
Machine Heart
Machine Heart
Sep 16, 2026 · Artificial Intelligence

Harness Evolution vs. Test-Time Scaling: Simple Retries Outperform Complex Self-Improvement

A study from AI2 and University of Washington finds that complex Harness Evolution for AI agents fails to consistently outperform simple test-time scaling methods like parallel sampling under equal compute budgets, and improvements rarely transfer to unseen tasks, questioning whether observed gains stem from genuine self-improvement or just extra attempts.

AI agentsAgent EvaluationBenchmarking
0 likes · 14 min read
Harness Evolution vs. Test-Time Scaling: Simple Retries Outperform Complex Self-Improvement
dbaplus Community
dbaplus Community
Sep 15, 2026 · Artificial Intelligence

Meituan's Agent Evaluation System: From Scoring to Infrastructure Capability

Meituan's Turing team shares a comprehensive framework for evaluating AI agents, covering multi-layered assessment (result, process, efficiency, risk), human-machine alignment via binary rubrics, seed test sets, expert knowledge integration, and the shift toward infrastructure for long-horizon agents with task-based evaluation harnesses.

Agent EvaluationEvaluation InfrastructureHuman-Machine Alignment
0 likes · 33 min read
Meituan's Agent Evaluation System: From Scoring to Infrastructure Capability
dbaplus Community
dbaplus Community
Sep 9, 2026 · Artificial Intelligence

Production-Grade Enterprise Agents: Unifying Harness, Skills & Virtual File Systems

This article details a production-grade architecture for enterprise AI agents, combining a unified harness for execution control, federated skills for domain expertise, and a virtual file system for long-task context management, drawing on Stripe's Kai platform and Deep Agents framework to address governance, security, and scalability challenges.

AI GovernanceAgent EvaluationAgent Harness
0 likes · 37 min read
Production-Grade Enterprise Agents: Unifying Harness, Skills & Virtual File Systems
Alibaba Cloud Native
Alibaba Cloud Native
Sep 4, 2026 · Artificial Intelligence

Close the Loop Finale: Agent Evaluation, SkillOps & LangChain Engineering Practices

The Shanghai finale of the Agent Observability and Optimization Close the Loop tour featured three technical sessions on building verifiable Agent optimization loops, managing Skills as strategic assets with SkillOps, and applying LangChain's Agent Engineering lifecycle from prototype to production, plus a hands-on workshop using Qoder and PawBench to demonstrate end-to-end evaluation.

Agent EvaluationAgent optimizationAgentLoop
0 likes · 9 min read
Close the Loop Finale: Agent Evaluation, SkillOps & LangChain Engineering Practices
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 26, 2026 · Artificial Intelligence

Can AI Do Independent Research? ASI‑Bench Measures Scientific Autonomy

ASI‑Bench, developed by Tsinghua and leading institutions, is a benchmark that evaluates AI’s scientific autonomy by progressively reducing method guidance across four levels, revealing that current models lose up to half their scientific score without detailed instructions, highlighting the gap to true independent research.

AI autonomyASI-BenchAgent Evaluation
0 likes · 14 min read
Can AI Do Independent Research? ASI‑Bench Measures Scientific Autonomy
Huolala Tech
Huolala Tech
Aug 12, 2026 · Artificial Intelligence

Redefining Agent Evaluation: An Offline LLM-as-a-Judge Framework and Practice

The article presents a quantitative offline evaluation system for AI agents in real outbound-call scenarios, combining reference‑based scoring with pairwise GSB methods, addressing regression and optimization, mitigating systematic bias through judge model selection and majority‑vote adjustments, and delivering an objective, high‑efficiency benchmark.

Agent EvaluationBias MitigationGSB
0 likes · 16 min read
Redefining Agent Evaluation: An Offline LLM-as-a-Judge Framework and Practice
Meituan Technology Team
Meituan Technology Team
Aug 6, 2026 · Artificial Intelligence

A Deep Dive into Agent Evaluation: From Basics to Advanced Practices

This article explains why evaluating AI agents requires more than answer correctness, outlines a four‑layer evaluation framework (result, process, efficiency, risk), compares short‑ and long‑horizon agents, and presents a practical methodology that combines objective and subjective metrics, rubric binary‑ization, case management, and infrastructure requirements for scalable, repeatable agent testing.

AI AgentAgent EvaluationLong-Horizon Agents
0 likes · 25 min read
A Deep Dive into Agent Evaluation: From Basics to Advanced Practices
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jun 24, 2026 · Artificial Intelligence

Can Agents Truly Self‑Evolve? GDPevo Benchmark That No Agent Can Cheat

The article introduces GDPevo, the first open‑source benchmark that quantifies self‑evolution in agents by generating 120 real‑world enterprise tasks, using rule‑hybrid question creation and deterministic scoring, and shows that self‑evolving agents improve accuracy by 17‑22% while reducing token consumption.

AI benchmarkAgent EvaluationContinual Learning
0 likes · 12 min read
Can Agents Truly Self‑Evolve? GDPevo Benchmark That No Agent Can Cheat
AI Engineering
AI Engineering
Jun 1, 2026 · Artificial Intelligence

Why Do Most Agent Projects Fail Before Launch? LangChain’s Solution

The article explains why many AI Agent projects collapse before production due to non‑determinism, error propagation, and creative solutions, and presents LangChain’s Deep Agent evaluation framework—integrated with LangSmith, AWS Bedrock, and Pytest—to provide a reproducible, end‑to‑end testing and monitoring process.

AWS BedrockAgent EvaluationDeep Agent
0 likes · 9 min read
Why Do Most Agent Projects Fail Before Launch? LangChain’s Solution
PaperAgent
PaperAgent
May 25, 2026 · Artificial Intelligence

DeepSeek’s Harness: How Agent Harness Engineering Is Shaping the Next LLM Agent Era

The article surveys DeepSeek’s Harness initiative, presenting the Binding‑Constraint Thesis, three‑stage evolution from prompt to harness engineering, the ETCLOVG seven‑layer architecture, and concrete benchmark evidence that harness‑only improvements far outweigh model upgrades, while detailing security, observability, and governance considerations for reliable LLM agents.

AI architectureAgent EvaluationAgent Harness Engineering
0 likes · 12 min read
DeepSeek’s Harness: How Agent Harness Engineering Is Shaping the Next LLM Agent Era
ITPUB
ITPUB
May 16, 2026 · Artificial Intelligence

Managing AI‑Generated Code with an Agent‑Based Evaluation Framework: Lessons from Refactoring 310 K Lines

When over 90% of a codebase is produced by AI, the authors show how a unified "people‑align → human‑machine‑align" approach, driven by evaluation agents, transforms technical debt into incremental business work, enabling continuous refactoring, AI‑friendly standards, and a sustainable engineering environment.

AI GovernanceAI codingAgent Evaluation
0 likes · 21 min read
Managing AI‑Generated Code with an Agent‑Based Evaluation Framework: Lessons from Refactoring 310 K Lines
Meituan Technology Team
Meituan Technology Team
May 7, 2026 · R&D Management

Managing AI‑Generated Code with Agent‑Based Evaluation: Refactoring 310K Lines of Code

When over 90% of a codebase is produced by AI, system quality hinges on constraining AI rather than speed, and this article details how a team used an agent‑based evaluation framework, unified standards, and incremental refactoring to turn 310,000 lines of AI‑written code into a maintainable, low‑debt system.

AI GovernanceAI codingAgent Evaluation
0 likes · 21 min read
Managing AI‑Generated Code with Agent‑Based Evaluation: Refactoring 310K Lines of Code
AntData
AntData
Apr 28, 2026 · Artificial Intelligence

Iterative Agent Evaluation Skill: Automating Bad‑Case Diagnosis with AI Pre‑Annotation

The article presents an end‑to‑end, eight‑phase automated evaluation pipeline for large‑model agents that replaces manual bad‑case inspection with AI‑assisted pre‑annotation, cutting analysis time from a full‑day to about 30 minutes and achieving over 90 % efficiency gain while enabling iterative knowledge‑base refinement.

AI pre‑annotationAgent EvaluationAutomated Pipeline
0 likes · 20 min read
Iterative Agent Evaluation Skill: Automating Bad‑Case Diagnosis with AI Pre‑Annotation
Alibaba Cloud Developer
Alibaba Cloud Developer
Mar 27, 2026 · Artificial Intelligence

How OpenClaw Empowers a Self‑Evolving Bank Manager Assistant

This article details a three‑day deep dive into OpenClaw, demonstrating how a self‑iterating AI assistant for bank relationship managers can be built, validated, and refined through autonomous agent communication, scheduled tasks, and memory‑driven reflection.

AI agentsAgent EvaluationOpenClaw
0 likes · 20 min read
How OpenClaw Empowers a Self‑Evolving Bank Manager Assistant
PaperAgent
PaperAgent
Dec 23, 2025 · Artificial Intelligence

CATArena: A Competitive Benchmark That Turns Agent Scoring into Evolutionary Learning

CATArena introduces a tournament‑style evaluation framework where AI agents iteratively code, compete, and improve across classic board games, using three‑dimensional quantitative scores to measure strategy programming, global learning, and generalization, and reveals how different LLM‑based agents learn and adapt over multiple rounds.

AI benchmarkAgent EvaluationCATArena
0 likes · 8 min read
CATArena: A Competitive Benchmark That Turns Agent Scoring into Evolutionary Learning
Amazon Cloud Developers
Amazon Cloud Developers
Dec 23, 2025 · Artificial Intelligence

Evaluating Agent Quality: A Practical Guide for Agentic AI

This article explains why evaluating AI agents is essential, outlines a multi‑dimensional metric system covering performance, safety, cost and bias, describes common evaluation frameworks such as AgentBoard, AgentBench and τ‑bench, and provides step‑by‑step instructions, example datasets and code for building a robust agent assessment pipeline.

AI agentsAgent EvaluationBenchmarking
0 likes · 35 min read
Evaluating Agent Quality: A Practical Guide for Agentic AI
Fun with Large Models
Fun with Large Models
Aug 20, 2025 · Artificial Intelligence

DeepSeek V3.1 Review: 128K Context, Knowledge, Programming & Agent Skills Near Claude 4

DeepSeek V3.1, released on August 19, expands context length to 128 K tokens and updates its knowledge base to July 2024, and the author’s benchmarks show its programming and agent capabilities now rival Claude 4, with detailed prompt examples, code generation demos, and performance comparisons.

Agent EvaluationClaude 4Context Length
0 likes · 9 min read
DeepSeek V3.1 Review: 128K Context, Knowledge, Programming & Agent Skills Near Claude 4
DataFunTalk
DataFunTalk
Jul 14, 2025 · Artificial Intelligence

Can Kimi K2 Beat Claude and Gemini in Coding and Agent Tasks?

This in‑depth review examines Kimi K2’s new focus on agent and coding abilities, comparing its performance on 3D HTML generation, code generation, and real‑world agent tasks against Claude 4 and Gemini 2.5, while also evaluating cost, openness, and practical usability for developers.

AI codingAgent EvaluationKimi K2
0 likes · 15 min read
Can Kimi K2 Beat Claude and Gemini in Coding and Agent Tasks?