Tagged articles

long-horizon evaluation

2 articles · Page 1 of 1
Machine Heart
Machine Heart
Sep 15, 2026 · Artificial Intelligence

WorldRoamBench: Benchmarking Long-Horizon Stability in Interactive World Models

Amap, Nanjing University, Tsinghua, and Peking University release WorldRoamBench, a benchmark with 1000+ long-horizon roaming samples evaluating 12 interactive world models on action following, visual stability, physics adherence, and memory retention, revealing that even top models like Genie 3 show significant weaknesses in specific dimensions.

Genie 3WorldRoamBenchaction following
0 likes · 9 min read
WorldRoamBench: Benchmarking Long-Horizon Stability in Interactive World Models
Machine Heart
Machine Heart
Jul 23, 2026 · Artificial Intelligence

Workflow Gym: Moving Beyond Simulated Tests to Bridge the Gap for GUI Agents

Workflow Gym introduces a realistic, long‑horizon benchmark covering 56 professional applications and 338 real‑world workflows, revealing that top GUI agents like Gemini 3.1 Pro achieve only about 30% single‑run success and exposing key failure modes such as consistency breaks and lack of domain knowledge.

AI performanceBenchmarkGUI agents
0 likes · 14 min read
Workflow Gym: Moving Beyond Simulated Tests to Bridge the Gap for GUI Agents