Statistics · AI Crawler Monitor

How many German websites block AI crawlers?

A daily robots.txt measurement across Germany’s top 1000 websites: which AI bots get locked out, how often — and what the 20 biggest sites do. All numbers are free to cite.

As of 30 September 2026 · Panel: Deutschland (CrUX Top-1000), 868 sites reachable

30.2 %of the top websites lock out GPTBot via robots.txt (incl. blanket blocks for all bots)
15 / 20of the 20 best-known sites block GPTBot
868websites in the daily measurement panel
12AI crawlers tracked

Block rate per AI crawler

Share of reachable panel sites that lock out each bot via robots.txt, explicitly or with a blanket block (User-agent: *):

CCBot · Common Crawl · Dataset30.8 %
GPTBot · OpenAI · Training30.2 %
Bytespider · ByteDance · Training28.5 %
ClaudeBot · Anthropic · Training26.6 %
meta-externalagent · Meta · Training23.8 %
Google-Extended · Google · Gemini training23.7 %
Applebot-Extended · Apple · Training23.4 %
anthropic-ai · Anthropic · Training (legacy)21.4 %
PerplexityBot · Perplexity · Search20.6 %
Amazonbot · Amazon · Assistant20.4 %
ChatGPT-User · OpenAI · On demand18.4 %
OAI-SearchBot · OpenAI · Search14.4 %

Method & context

Every day the robots.txt of Germany’s top 1000 websites (CrUX panel) is measured; for 868 reachable sites the rules for 12 known AI crawlers are evaluated. A subset is additionally probed for server-side blocks (e.g. 403 for bot user agents).

“Blocked” here means: robots.txt locks out the bot explicitly or with a blanket block for all bots (User-agent: *). The interactive monitor shows server-side blocks separately; they are not part of these rates. Sites without a rule count as not blocking, and robots.txt is a request, not a technical barrier.

Tracked since 3 July 2026. The figures on this page come from the measurement dated above; the interactive AI Crawler Monitor shows the current day’s measurement.