Harness Evolution vs. Test-Time Scaling: Simple Retries Outperform Complex Self-Improvement
A study from AI2 and University of Washington finds that complex Harness Evolution for AI agents fails to consistently outperform simple test-time scaling methods like parallel sampling under equal compute budgets, and improvements rarely transfer to unseen tasks, questioning whether observed gains stem from genuine self-improvement or just extra attempts.
