Cloudflare lets sites block AI training crawlers while staying searchable
djfergus · hn · 2026-09-16
Cloudflare's new blog post introduces "accountable mixed-use crawlers," addressing the long-standing dilemma where robots.txt forces an all-or-nothing choice: allow crawlers for search traffic or block them entirely.
The new approach lets site owners stay discoverable in search while explicitly disallowing the same crawlers from using content for AI training, adding accountability requirements for mixed-use bots. Discussion is ongoing on Hacker News about feasibility and crawler compliance.
More from Infra
- SentencePiece Lite ships: 50KB binary, 20-30x faster tokenization for edge devices — heiga_zen · 2026-09-16
- Leaked 10,000-word Huawei memo: become the NVIDIA for any LLM, pivot around Ascend 950 — pstAsiatech · 2026-09-16
- Running Qwen 27B and DeepSeek v4 Flash together on one heterogeneous machine — samsja19 · 2026-09-16
- JPMorgan sees 25M+ GPU/ASIC shipments by 2028, ASICs dominate — a 'narrative violation' — bookwormengr · 2026-09-16
- iamtrask: The Endgame Is a Trust Web of Personal LLM Servers, Not One AGI — iamtrask · 2026-09-16
- Program-as-Weights: 0.6B interpreter matches Qwen3-32B prompting with 1/50 the memory — yuntiandeng · 2026-09-16