Optical Illusions as a New Attack Surface for AI Safety

plopesresearch · x · 2026-07-13

This post highlights an AI security question: Could optical illusions become an attack surface for bypassing AI systems, much like zero-day vulnerabilities?

The core argument is that while known illusions are relatively easy to spot, the real challenge is that different illusions exploit different perceptual biases, making it difficult to build defenses that generalize against "unknown illusions." The author cites an example of a "font only humans can read," which could not be correctly interpreted by Fable and GPT 5.6 Sol Ultra.

These novel illusions and visual encodings could allow information to pass through AI systems undetected while remaining perfectly legible to humans, effectively bypassing various AI safeguards.

Original post →

More from Safety

Safety channel →