Lasso finds AI text watermarks change how LLM agents act
CackleRooster · reddit · 2026-09-22
Lasso Security reports that LLM text watermarks—meant to answer whether a machine wrote something—also alter how the model behaves, with implications for agents built on watermarked models. Watermarking may not be a side-effect-free signal.
More from Models
- Unverified DataBench Charts Fuel Rumors of OpenAI's Internal Model 'Luna' Ahead of GPT-6 — almmaasoglu · 2026-09-22
- OpenAI's 24-day-old internal model reportedly solved 100+ open math problems — IgorCarron · 2026-09-22
- Paradigm unveils Limite 1B 'Violetto', a 1B model built for high-frequency mathematical intelligence — tensorqt · 2026-09-22
- Speculation: Grok Pro line is an extension of Flash line, mxfp4 QAT likely speeds RL rollouts — stochasticchasm · 2026-09-22
- Adam-to-Muon mid-training switch sparks debate, seemingly contradicting Moonlight paper — stochasticchasm · 2026-09-22
- Diffusion LMs were 5-10x faster but never loved — timing, not speed, wins — eliza_luth · 2026-09-22