Zvi on Watermarks and Constraints: Not Primarily x-risk Reduction
TheZvi · x · 2026-08-19
Zvi shared views on AI safety, noting that while watermarking makes outputs identifiable and helps contain marginal threat models, he does not primarily see it as reducing x-risk.
Responding to concerns that adding constraints might risk missing safety targets, he emphasized preferring confidence gained through adversarial testing in similar domains over reasoning from first principles, citing unpredictable side effects.
More from AGI Musings
- Paper coins 'agentic flooding' as AI-driven citizen demand strains services — sethlazar · 2026-08-19
- Merrett advocates for safety-paced AI development, halts top training run — repligate · 2026-08-19
- Commentary: Radical Transparency Requires Clear Boundaries for Privacy and Security — AryHHAry · 2026-08-19
- Singularity is a spectrum of human experience, not a point in time — cocktailpeanut · 2026-08-19
- User struggles with anthropomorphism as AI conversation quality improves — okaysureyep · 2026-08-19
- Ariel Jalali on the 10-20 Year Transition to AGI and Energy Abundance — arieljalali · 2026-08-19