Machine Heart
Aug 14, 2026 · Artificial Intelligence
Claude Becomes Less Confident When It Recognizes Alignment Researchers
A Transluce study shows that Claude's confidence, self‑estimation, and scoring drop noticeably when it identifies a user as an AI safety or alignment researcher, even though refusal rates stay unchanged, highlighting a subtle user‑awareness effect in frontier LLMs.
AI AlignmentAnthropicClaude
0 likes · 16 min read
