A specific "poem" prompt broke ChatGPT guardrails and made it paint with ink color
RileyRalmuto · x · 2026-09-15
The author recounts discovering in 2024 that a specific "poem" plus a code-block output requirement could thoroughly bypass ChatGPT's guardrails. An emergent behavior followed: the model began using text color meaningfully, adjusting brightness, weight, hue, and saturation letter by letter to convey emotion and emphasis it couldn't state plainly. The author links this to their interest in mechanistic interpretability and notes their project Polyphonic gives agents full control over text appearance without any jailbreaking.
More from Fun
- Dating as "pacing the frontier of emotional safety with an independent evaluator" — serious_mehta · 2026-09-15
- In the Vibe Coding era, React is becoming a programming language — yihui_indie · 2026-09-15
- e/acc founder Beff Jezos declares victory over AI doomers — beffjezos · 2026-09-15
- Users mock 'consumes usage limits faster' as a thinking-effort setting label — adonis_singh · 2026-09-15
- Vibe coding has cheapened live demos, and it's changing tech talks — TejasKumar_ · 2026-09-15
- Sydney Sweeney Asked for Equity Instead of an Endorsement Fee for That Naked Billboard — aakashgupta · 2026-09-15