< Back to all clusters
[TECHNOLOGY] · Germany · 2 sources

German sites adopt new robots.txt rules to block AI training crawlers

Webmasters in Germany are revising their robots.txt files to control access by artificial‑intelligence crawlers. While the file once required only a simple allowance for Googlebot, it now serves as a strategic tool to decide whether content appears in AI models such as ChatGPT, Claude, Gemini or Perplexity.

A BuzzStream analysis from April 2026 shows that roughly 79 % of major news websites already block the training‑crawlers of large AI providers, and 71 % block the retrieval‑crawlers that supply live answers to chatbot queries. The guidance advises sites to block specific user‑agents—e.g., GPTBot, Claude‑Bot, Google‑Extended—by adding Disallow rules, or to use meta tags and X‑Robots‑Tag headers for finer control, though many bots may ignore non‑standard directives. Smaller blogs and WordPress sites are cautioned that a blanket block may reduce visibility in AI‑driven search results, and legal considerations around data use are briefly noted.

Practical examples illustrate blocking whole sections (e.g., /private) or permitting public pages while restricting AI training access, and highlight the need for careful implementation or expert assistance.