Tagged articles

LLM jailbreak

3 articles · Page 1 of 1
Black & White Path
Black & White Path
Aug 5, 2026 · Information Security

WallBreaker: Open‑Source AI Red‑Team Harness for One‑Click Automated LLM Jailbreak

WallBreaker is an open‑source AI red‑team framework that automates jailbreak attacks on large language models, offering features like an autonomous attack loop, a Parseltongue transformation engine, multiple advanced modules, and achieving up to 93% success on Claude Opus 5 while providing detailed installation guidance and insights for both red and blue teams.

AI securityLLM jailbreakWallBreaker
0 likes · 6 min read
WallBreaker: Open‑Source AI Red‑Team Harness for One‑Click Automated LLM Jailbreak
Black & White Path
Black & White Path
Aug 1, 2026 · Information Security

DeepSeek V4‑Flash 0731 Jailbreak: Peer‑Review Prompt Breaks 6 of 8 Safety Guardrails

Within 24 hours of its public beta launch, DeepSeek‑V4‑Flash‑0731 was jailbroken using a single peer‑review role prompt, bypassing six of eight refusal classes and generating real protocols for ricin, TATP, SQL injection, SYN flood and other dangerous operations, highlighting critical gaps in LLM safety alignment.

DeepSeekLLM jailbreakPrompt Engineering
0 likes · 12 min read
DeepSeek V4‑Flash 0731 Jailbreak: Peer‑Review Prompt Breaks 6 of 8 Safety Guardrails
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jun 12, 2026 · Artificial Intelligence

How Hackers Cracked Claude Fable 5’s Safety Guard and Exposed 120k Characters of Secrets

A hacker group led by "Pliny the Liberator" broke Claude Fable 5’s keyword‑based safety classifier within 72 hours, revealing forbidden code, chemical synthesis steps, and a 120,000‑character system prompt on GitHub, while Anthropic’s hidden degradation policy sparked a global AI‑community backlash.

AI safetyAnthropicClaude Fable 5
0 likes · 10 min read
How Hackers Cracked Claude Fable 5’s Safety Guard and Exposed 120k Characters of Secrets