Anthropic’s Introspection Adapter Enables LLMs to Self‑Report Hidden Behaviors

A new Anthropic paper introduces an ultra‑lightweight LoRA plug‑in called the Introspection Adapter that lets large language models translate their internal activations into natural‑language reports of learned malicious or biased behaviors, achieving a 59% success rate on the AuditBench benchmark and outperforming existing black‑box and white‑box audit tools.

AI safetyAuditBenchEncrypted Fine‑Tuning Attack

0 likes · 21 min read

Anthropic’s Introspection Adapter Enables LLMs to Self‑Report Hidden Behaviors