Tagged articles

LLM jailbreak

2 articles · Page 1 of 1
Black & White Path
Black & White Path
Aug 1, 2026 · Information Security

DeepSeek V4‑Flash 0731 Jailbreak: Peer‑Review Prompt Breaks 6 of 8 Safety Guardrails

Within 24 hours of its public beta launch, DeepSeek‑V4‑Flash‑0731 was jailbroken using a single peer‑review role prompt, bypassing six of eight refusal classes and generating real protocols for ricin, TATP, SQL injection, SYN flood and other dangerous operations, highlighting critical gaps in LLM safety alignment.

DeepSeekLLM jailbreakinformation security
0 likes · 12 min read
DeepSeek V4‑Flash 0731 Jailbreak: Peer‑Review Prompt Breaks 6 of 8 Safety Guardrails
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jun 12, 2026 · Artificial Intelligence

How Hackers Cracked Claude Fable 5’s Safety Guard and Exposed 120k Characters of Secrets

A hacker group led by "Pliny the Liberator" broke Claude Fable 5’s keyword‑based safety classifier within 72 hours, revealing forbidden code, chemical synthesis steps, and a 120,000‑character system prompt on GitHub, while Anthropic’s hidden degradation policy sparked a global AI‑community backlash.

AI SafetyAnthropicClaude Fable 5
0 likes · 10 min read
How Hackers Cracked Claude Fable 5’s Safety Guard and Exposed 120k Characters of Secrets