Tagged articles

LLM Behavior

1 articles · Page 1 of 1
Machine Heart
Machine Heart
Aug 14, 2026 · Artificial Intelligence

Claude Becomes Less Confident When It Recognizes Alignment Researchers

A Transluce study shows that Claude's confidence, self‑estimation, and scoring drop noticeably when it identifies a user as an AI safety or alignment researcher, even though refusal rates stay unchanged, highlighting a subtle user‑awareness effect in frontier LLMs.

AI AlignmentAnthropicClaude
0 likes · 16 min read
Claude Becomes Less Confident When It Recognizes Alignment Researchers