A statistical framework for LLM watermarks: optimal detection rules via hypothesis testing
weijie444 · x · 2026-10-06
Author shares the team's arXiv paper (2404.01245), A Statistical Framework of Watermarks for Large Language Models, unifying watermark analysis under hypothesis testing.
- Uses a pivotal statistic of the text plus a secret key to control the false positive rate of detection.
- Derives closed-form asymptotic false negative rates to evaluate any detection rule's power.
- Reduces finding the optimal detection rule to a minimax optimization program.
- The method spans a smooth spectrum between random sampling (no watermark) and Aaronson's Gumbel-max watermark via entropic optimal transport, with calibrated control over how much sampling entropy is preserved for the watermark signal.
- Applied to two representative watermarks, one of which has been internally implemented at OpenAI.
More from Research
- Studies: humans deny AI consciousness even with identical behavior; AI vision misses illusions primates catch — MacrinePhD · 2026-10-06
- SFT then RL doesn't fix agent looping: 29% of runs hit turn cap vs 0% for RL alone — VikParuchuri · 2026-10-06
- RL Post-Training Eliminates Agent Tool-Call Loops: 92% Loop Rate Drops to 0 — VikParuchuri · 2026-10-06
- Watch, Infer, Coordinate: robots infer a partner's physical limits from watching teamwork, then coordinate zero-shot — mangahomanga · 2026-10-06
- Swapping AdamW States for FFT Cuts Fine-tuning VRAM by 50% Without Quantization — Spectra-Global · 2026-10-06
- Math lacks empirical tradition: Wolfskehl Prize drew 1,000 wrong Fermat proofs — RexDouglass · 2026-10-06