Tagged articles

long‑horizon evaluation

1 articles · Page 1 of 1
Machine Heart
Machine Heart
Jul 23, 2026 · Artificial Intelligence

Workflow Gym: Moving Beyond Simulated Tests to Bridge the Gap for GUI Agents

Workflow Gym introduces a realistic, long‑horizon benchmark covering 56 professional applications and 338 real‑world workflows, revealing that top GUI agents like Gemini 3.1 Pro achieve only about 30% single‑run success and exposing key failure modes such as consistency breaks and lack of domain knowledge.

AI performanceGUI agentsWorkflow Gym
0 likes · 14 min read
Workflow Gym: Moving Beyond Simulated Tests to Bridge the Gap for GUI Agents