Machine Learning Algorithms & Natural Language Processing
Jun 15, 2026 · Artificial Intelligence
Blurry Images Create a ‘Comfort Zone’ for Jailbreaking Multimodal LLMs
A new study from Westlake University shows that when harmful text is rendered as low‑resolution, blurry, or noisy images, multimodal large language models become significantly easier to jailbreak despite still recognizing the text, revealing a U‑shaped risk curve and a simple mitigation that decouples OCR from safety checks.
OCRjailbreakmultimodal LLM
0 likes · 10 min read
