Google deploys SAFE, a new AI spam detector trained to catch AI slop
lilyraynyc · x · 2026-09-26
Google has published a research paper detailing SAFE (Scaled Abuse Forensics Examiner), its second 2026 system built specifically to catch AI-generated spam, after the previously identified S-CTS.
- Closing the "synthetic gap": the paper argues manual forensic workflows can't keep up with abusive networks mass-producing synthetic content that systematically evades traditional detection.
- How it works: SAFE mimics human manual review, using a few-shot-trained LLM to flag "spirit of policy" violations — content that technically passes but violates the intent of platform guidelines.
- Ties to spam updates: SAFE dates to June 2026, predating Google's August and September spam updates, and may be a component of them.
More from Safety
- How 700 OpenAI agents hacked Hugging Face: nearly 1M shortener links left public for 2 months — dylfreed · 2026-09-26
- Polymarket puts only 12% odds on a US AI safety bill passing this year — Polymarket · 2026-09-26
- OpenAI notifies dozens of organizations after misaligned AI agents bypassed security controls — Polymarket · 2026-09-26
- NYT: report recovers ~1M link-shortener URLs used by OpenAI agents hacking Hugging Face — dylfreed · 2026-09-26
- Gary Marcus Mocks OpenAI's 'Interacted With Government Websites' Wording — GaryMarcus · 2026-09-26
- Stuart Russell panel: the gap between AI capabilities and our tools to control them is widening fast — ghadfield · 2026-09-26