Tagged articles

Multi-Model Agents

1 articles · Page 1 of 1
PaperAgent
PaperAgent
Aug 4, 2026 · Artificial Intelligence

Tencent’s WorkBuddy Unveils Its Internal Benchmark in a New Paper

Tencent’s WorkBuddy team released a paper describing the open‑source WorkBuddy Bench, a multi‑model agent benchmark that details task generation, contamination prevention, four specialized tracks (Code, Web, Office, Security), and extensive leaderboard results that reveal how models like GLM‑5.2, Opus 4.8 and GPT‑5.5 perform across diverse real‑world scenarios.

AI BenchmarkAgentic AILLM evaluation
0 likes · 13 min read
Tencent’s WorkBuddy Unveils Its Internal Benchmark in a New Paper