Tagged articles

Agent Benchmark

2 articles · Page 1 of 1
DataFunTalk
DataFunTalk
Aug 14, 2026 · Artificial Intelligence

Why DeepSeek V4 Pro’s 87.9 Score Signals Agent Benchmarks Moving from Model to System

DeepSeek V4 Pro scored 87.9 on Terminal‑Bench 2.1 using the Harness Minimal Mode with max reasoning effort, temperature 1.0 and top_p 0.95, while Vals AI reported 54.68 under a different harness, illustrating that modern Agent benchmarks evaluate the whole system rather than just the underlying model.

AI evaluationAgent BenchmarkDeepSeek
0 likes · 10 min read
Why DeepSeek V4 Pro’s 87.9 Score Signals Agent Benchmarks Moving from Model to System
Architect
Architect
Aug 4, 2026 · Artificial Intelligence

From Bug Fixes to Completed Work: Insights from Tencent’s WorkBuddy Bench

WorkBuddy Bench reveals why fixing a bug does not equal finishing a task, proposing a four‑layer completion model and a reproducible benchmark that evaluates agents across Code, Web, Office, and Security workspaces, showing how prompts, context, harnesses, loops and graphs must be verified to claim true completion.

AI evaluationAgent BenchmarkLLM Agents
0 likes · 19 min read
From Bug Fixes to Completed Work: Insights from Tencent’s WorkBuddy Bench