Why Max Reasoning Backfires: Opus 5.5 Medium Outperforms Max in Terminal-Bench 4.0
Terminal-Bench 4.0 reveals that excessive reasoning (max) reduces accuracy for top models like Opus 5.5 and GPT-6 Astra due to overthinking traps—agent drift, over-engineering, context dilution, and self-doubt—while high-end models at medium reasoning outperform mid-tier models at max reasoning at lower cost.
