Cloudflare's new setting lets sites refuse AI training while keeping Google indexing
shashib · x · 2026-09-24
On September 15 Cloudflare launched a "Disallow AI Training" dashboard setting that lets sites refuse AI model crawlers while still allowing Google, Bing, and Apple search indexing, addressing the long-standing blur between search and training permissions in robots.txt.
- Cloudflare data: 36.6% of verified crawler traffic is mixed-use; <1% of sites block search crawlers entirely, and 17% restrict training.
- Guidance: use Disallow AI Training to keep search but refuse training; use Block only if losing Googlebot is acceptable; Allow permits training crawlers.
- Bing's robots.txt training-preference support is targeted for 2027.
The piece also warns of the opposite failure mode: over-blocking can make pages invisible to AI assistants, so a day's work on a press release may never reach a reader.
More from Infra
- AI Gateway explained in 2 minutes: routing, auth, caching and observability — _jaydeepkarale · 2026-09-24
- Jev turns past LLM responses into a semantic cache to cut agent inference costs — MikkoH · 2026-09-24
- AI Compute Firm Fluidstack Opens 8 Roles Around 'Decision Engineering' for Gigawatt-Scale Infrastructure — MxMnr · 2026-09-24
- Multi-GPU monitor plugin released for DeepSeek Harness, seeking Windows testers — paulqq · 2026-09-24
- Yoniq Compute open-sources inference recipes for SGLang/vLLM, validated on up to 8x H200 — TheZachMueller · 2026-09-24
- Microsoft to invest over $10B across UAE, Saudi Arabia, Qatar and Kuwait by 2030 — Polymarket · 2026-09-24