Tagged articles

Harness Evolution

2 articles · Page 1 of 1
Machine Heart
Machine Heart
Sep 17, 2026 · Artificial Intelligence

Beyond Prompts: Harness-Policy Co-Evolution for Agent Safety by Shanghai AI Lab

Shanghai AI Lab and university collaborators propose SHE and SafeEvolve, two frameworks that evolve agent safety by learning from execution trajectories: SHE updates a modular safety harness via trajectory-driven evolution, while SafeEvolve distills verified harness experience into the policy model through SFT and RL, reducing attack success rates on benchmarks.

Agent SafetyHarness EvolutionLLM Agents
0 likes · 12 min read
Beyond Prompts: Harness-Policy Co-Evolution for Agent Safety by Shanghai AI Lab
Machine Heart
Machine Heart
Sep 16, 2026 · Artificial Intelligence

Harness Evolution vs. Test-Time Scaling: Simple Retries Outperform Complex Self-Improvement

A study from AI2 and University of Washington finds that complex Harness Evolution for AI agents fails to consistently outperform simple test-time scaling methods like parallel sampling under equal compute budgets, and improvements rarely transfer to unseen tasks, questioning whether observed gains stem from genuine self-improvement or just extra attempts.

AI agentsAgent EvaluationBenchmarking
0 likes · 14 min read
Harness Evolution vs. Test-Time Scaling: Simple Retries Outperform Complex Self-Improvement