PaperAgent
Aug 4, 2026 · Artificial Intelligence
Tencent’s WorkBuddy Unveils Its Internal Benchmark in a New Paper
Tencent’s WorkBuddy team released a paper describing the open‑source WorkBuddy Bench, a multi‑model agent benchmark that details task generation, contamination prevention, four specialized tracks (Code, Web, Office, Security), and extensive leaderboard results that reveal how models like GLM‑5.2, Opus 4.8 and GPT‑5.5 perform across diverse real‑world scenarios.
AI BenchmarkAgentic AILLM evaluation
0 likes · 13 min read
