Muon Alone Can't Remove Watermarks; Needs Layered Learning Rates
willdepue · x · 2026-07-14
Discussion around experimental findings regarding watermark removal/perturbation models:
- Using Muon alone appears insufficient to fully erase the watermark.
- Combining Muon with Modula's layered learning rates yields results much closer to complete "erasure."
- However, the author notes that Modula performs poorly on non-toy problems, making this combination unreliable for real-world tasks.
Overall, this suggests that while certain optimization methods are more effective at disrupting model parameters/traces, we are still far from a stable, universal watermark removal solution.
More from Research
- Project APE finds verifier reliability drops when papers contain multiple errors — soumitrashukla9 · 2026-07-22
- Project APE says verifier costs fell about 90x in a year as Chinese open models lead — soumitrashukla9 · 2026-07-22
- OpenAI-linked paper says capability RL can make models more reward-seeking — MariusHobbhahn · 2026-07-22
- Project APE builds its verifier benchmark from 100 AI-written papers with injected errors — soumitrashukla9 · 2026-07-22
- Paper proposes a CRED taxonomy and benchmark to measure research-error detectors — soumitrashukla9 · 2026-07-22
- OpenAI says long-horizon models need safety and alignment checks across full action sequences — rhiever · 2026-07-22