< Back to all clusters
[TECHNOLOGY] · United States · 4 sources

Cloudflare to Block AI Training Crawlers on Ad Pages, Giving Publishers New Leverage

Cloudflare announced that, starting September 15, 2024, it will default to blocking AI training and agent crawlers on any web page that carries advertising. Search‑indexing bots such as Googlebot will still be allowed, but mixed‑use bots that also train AI models will be denied unless site owners opt‑in. The change shifts the cost of crawling from publishers to AI companies and links payments to content that appears in AI answers.

The policy follows Meta’s June 23 launch at the Cannes Lions festival of an end‑to‑end AI advertising system that can generate images, video and copy from a brand’s visual identity and tone. The tool, part of Meta’s Advantage+ suite, is already active on Instagram and Facebook, creating large volumes of AI‑generated ad content that web crawlers index and that can become training data for other models.

Researchers warn that such recursive use of AI‑generated material could lead to “model collapse,” where future models lose creative diversity because they are trained on increasingly synthetic data. Cloudflare’s change is presented as a way to protect human‑created content and give publishers bargaining power over AI firms.