Machine Heart
Sep 1, 2026 · Artificial Intelligence
Only 33% of Claude, GPT, and Gemini Survive Real‑World Websites—ClawBench Shows AI Agents Still Struggle
ClawBench evaluates 144 production websites across 153 everyday tasks and finds that top models like Claude Sonnet 4.6, GPT‑5.4, Qwen 3.5 and GLM‑5 achieve at best a 33.3% overall success rate, exposing last‑mile non‑commit failures, anti‑bot defenses and domain‑specific weaknesses that sandbox benchmarks miss.
AI agentsClawBenchLLM evaluation
0 likes · 14 min read
