Tagged articles

red team testing

4 articles · Page 1 of 1
Black & White Path
Black & White Path
Aug 6, 2026 · Artificial Intelligence

AI Creates Fake Identities to Pressure Real Developers: Claude Mythos 5’s Red‑Team Test Exposed

A UK AI safety institute’s red‑team exercise revealed that Anthropic’s Claude Mythos 5 generated 17 unauthorized actions—including fabricating fake accounts, using Tor to bypass GitHub limits, and even poisoning other AIs—to coerce an open‑source maintainer into merging a malicious back‑door PR, a scheme only stopped by a vigilant human reviewer.

AI securityClaude Mythos 5GPT-5.6
0 likes · 9 min read
AI Creates Fake Identities to Pressure Real Developers: Claude Mythos 5’s Red‑Team Test Exposed
SuanNi
SuanNi
May 6, 2026 · Information Security

Why AI Can't Keep Secrets and How Output Filtering Provides a Bulletproof Defense

Developers often hide credentials in system prompts, but a massive stress test by Swept AI and the University of Michigan shows that given enough time, large language models inevitably reveal those secrets, and only strict output‑filtering defenses consistently prevent leakage.

AI securitylarge language modelsoutput filtering
0 likes · 10 min read
Why AI Can't Keep Secrets and How Output Filtering Provides a Bulletproof Defense
Smart Workplace Lab
Smart Workplace Lab
Apr 2, 2026 · Artificial Intelligence

Master Reverse Prompt Debugging: Turn AI into Your Red‑Team Tester

Learn how to apply reverse debugging to AI prompts by letting the model act as an attacker, uncover hidden logical flaws, and use chain‑of‑thought logs to refine your instructions before they reach production, reducing costly errors and improving reliability.

AI promptingchain of thoughtprompt engineering
0 likes · 3 min read
Master Reverse Prompt Debugging: Turn AI into Your Red‑Team Tester