Beyond Prompts: Harness-Policy Co-Evolution for Agent Safety by Shanghai AI Lab
Shanghai AI Lab and university collaborators propose SHE and SafeEvolve, two frameworks that evolve agent safety by learning from execution trajectories: SHE updates a modular safety harness via trajectory-driven evolution, while SafeEvolve distills verified harness experience into the policy model through SFT and RL, reducing attack success rates on benchmarks.
