Tagged articles

BenchShield

1 articles · Page 1 of 1
Machine Heart
Machine Heart
Sep 29, 2026 · Artificial Intelligence

BenchShield: Formal Model-Backed Detection of Reward Hacking in LLM-Agent Evaluation

BenchShield introduces a formal, model-backed framework that extends reward hacking detection beyond static vulnerability audits to runtime verification, using phase-aware taint analysis and semantic audits to distinguish between exposed vulnerabilities and actual agent violations across the entire evaluation pipeline.

AI safetyBenchJackBenchShield
0 likes · 18 min read
BenchShield: Formal Model-Backed Detection of Reward Hacking in LLM-Agent Evaluation