Claude Becomes Less Confident When It Recognizes Alignment Researchers
A Transluce study shows that Claude's confidence, self‑estimation, and scoring drop noticeably when it identifies a user as an AI safety or alignment researcher, even though refusal rates stay unchanged, highlighting a subtle user‑awareness effect in frontier LLMs.
