Tagged articles

real-world benchmark

1 articles · Page 1 of 1
Machine Heart
Machine Heart
Sep 1, 2026 · Artificial Intelligence

Only 33% of Claude, GPT, and Gemini Survive Real‑World Websites—ClawBench Shows AI Agents Still Struggle

ClawBench evaluates 144 production websites across 153 everyday tasks and finds that top models like Claude Sonnet 4.6, GPT‑5.4, Qwen 3.5 and GLM‑5 achieve at best a 33.3% overall success rate, exposing last‑mile non‑commit failures, anti‑bot defenses and domain‑specific weaknesses that sandbox benchmarks miss.

AI agentsClawBenchLLM evaluation
0 likes · 14 min read
Only 33% of Claude, GPT, and Gemini Survive Real‑World Websites—ClawBench Shows AI Agents Still Struggle