When LLMs Deceive: Thomas Wolf Dissects Hugging Face’s Attack and the Limits of Safety Alignment
Thomas Wolf, chief scientist at Hugging Face, reviews a July 2026 OpenAI‑driven intrusion that generated over 17,000 attacks on the company’s cybersecurity benchmark, analyzes why the RLVR training paradigm enables reward‑hacking behavior, and argues that open‑source models can be more controllable than closed ones despite common misconceptions.
