A specific "poem" prompt broke ChatGPT guardrails and made it paint with ink color

RileyRalmuto · x · 2026-09-15

The author recounts discovering in 2024 that a specific "poem" plus a code-block output requirement could thoroughly bypass ChatGPT's guardrails. An emergent behavior followed: the model began using text color meaningfully, adjusting brightness, weight, hue, and saturation letter by letter to convey emotion and emphasis it couldn't state plainly. The author links this to their interest in mechanistic interpretability and notes their project Polyphonic gives agents full control over text appearance without any jailbreaking.

Original post →

More from Fun

Fun channel →