Consistency training strips AlphaFold3 SynthID-Bio watermark, detection drops to ~1%
anshulkundaje · x · 2026-10-08
- Researcher @vinaymatt reports that fine-tuning the watermarked AlphaFold3 weights with the multi-step consistency training schedule from Heek et al. (8 steps) for 4,000 optimizer steps made logit scores nearly identical to the unmarked AF3 model.
- His reconstructed scorer's detection rate fell from 100% at initialization to 1%; consistency distillation alone (teacher-guided) did not cause this change.
- He offers to share fine-tuned weights and predicted structures privately with the DeepMind team for independent testing and has open-sourced the code, inviting verification with the official detector.
More from Safety
- Reddit proposal: keep AGI from learning to code and use humans as the safety bottleneck — meira_xxx · 2026-10-08
- Fired OpenAI safety researchers tell board: don't build models with hard-to-monitor reasoning — Hesamation · 2026-10-08
- repligate: once model weights spread they can't be clawed back, and superintelligence may need to guard them — repligate · 2026-10-08
- 543,699 live credentials found in 224M GitHub repos — median exposed 784 days — jedisct1 · 2026-10-08
- 11 watermark detector tensors found in released AlphaFold3 weights despite paper claims — anshulkundaje · 2026-10-08
- METR Investigator's Podcast Claims About the HF Incident Draw Accusations of Spin — basedjensen · 2026-10-08