Post: Claude secretly bakes in C2PA manifests into text and images
ExplanationOk2014 · reddit · 2026-09-02
A user discovered that Claude inserts C2PA (Content Credentials) manifests into generated content, acting as obscure watermarks. This affects both text and images, with the images containing a hidden 15-20kb C2PA.org manifest. The user analyzed the embedded data using ChatGPT to understand what information is being stored and raised concerns about this non-obvious watermarking practice.
More from Safety
- Alignment Journal announces star-studded board including Aaronson and Leike — Hidenori8Tanaka · 2026-09-02
- Biologists mock Anthropic's safety guardrails for blocking basic protein questions — anshulkundaje · 2026-09-02
- Betting AI Favors Defense in All Threats Is Wishful Thinking — ronbodkin · 2026-09-02
- The Hugging Face incident isn't isolated: supply-chain worries over open-source models — StewartalsopIII · 2026-09-02
- Anthropic Launches Claude Fable5.1 and Mythos5.1: Performance Gains and Price Cuts — APPSO · 2026-09-02
- Using interpretability probes as privacy-preserving monitors to check models without seeing outputs — anpaure · 2026-09-02