Should You Block AI Crawlers? A Decision Guide

Block or welcome AI is one trade-off: blocking buys control but forfeits citations. Decide per-crawler: block training, price retrieval, welcome citation.

Should you block AI crawlers, or welcome them?

Block the crawlers that never send anyone back and welcome the ones that do — the answer is per crawler and per content, not all-or-nothing. Blocking buys control over training and refuses no-referral crawlers but forfeits the citations and referrals that search and retrieval crawlers return. According to Cloudflare (2025), 80 % of the AI crawling it observed in the twelve months before its August 2025 report served training, 18 % search and 2 % user actions; only the last two send anyone back. Per-bot rationale for 41 crawlers and policy tokens is in this site's AI Crawler Registry (AGENTS WELCOME, 2026).

When is blocking AI crawlers the right call?

Block when a crawler takes without giving back and your content is your product. According to Cloudflare (2025), in July 2025 Anthropic's crawlers made 38,065 requests for every referral they sent back, OpenAI's 1,091 and Perplexity's 194, against 5.4 for Google. Blocking protects proprietary or paywalled content, refuses uncompensated training, and is the one lever a publisher fully controls; more than one million Cloudflare customers had chosen it by 1 July 2025, when Cloudflare gave new domains a permission-based default (Cloudflare, 2025). For a premium archive or licensable dataset, block-until-paid is coherent; pay-per-crawl and RSL turn it into revenue. One caveat: robots.txt rules "are not a form of access authorization" (IETF, 2022), so a hard block needs the CDN.

What does welcoming AI crawlers return?

Citations and referrals from the crawlers that feed AI answers — increasingly the only surface users see. According to Pew Research Center (2025), U.S. Google users clicked a traditional result in 8 % of visits when an AI summary appeared versus 15 % when none did, and ended their session on 26 % versus 16 % of such pages (68,879 searches, March 2025). Blocking those crawlers removes a site from the answers: sites opted out of OAI-SearchBot "will not be shown in ChatGPT search answers" (OpenAI, 2026). Of the 41 registry entries, 10 are search-index crawlers and 13 user-triggered fetchers against 9 training crawlers (AGENTS WELCOME, 2026). This site's stance is welcome — a position, not the only rational choice.

How do you decide per crawler: block, price or welcome?

Sort each crawler by what it returns: block uncompensated training, price high-value retrieval with pay-per-crawl or RSL, and welcome the citation crawlers that put your content in front of users. The table gives the default per crawler type; the decision turns on content type, referral value and licensing leverage. This site chooses welcome and proves it by configuration: robots.txt allows every named crawler with search=yes, ai-input=yes, ai-train=no, and /license.xml licenses training instead of refusing it (AGENTS WELCOME, 2026).

Default decision by crawler type with a registry example (AI Crawler Registry, updated 2026-07-06)
Crawler typeReturnsDefaultExample
Training crawlerno referral, no paymentblock, or price via RSLGPTBot, ClaudeBot, CCBot
Search-index crawlercitations, referralswelcome (Allow)OAI-SearchBot, PerplexityBot, Claude-SearchBot
User-triggered fetchera page for a real userwelcome (rate-limit if needed)ChatGPT-User, Claude-User
Crawler on premium contentpaid contentprice (HTTP 402 or RSL)paywalled archives

Blocking AI crawlers — frequently asked questions

Should I block AI crawlers?

It depends on the crawler and the content. Blocking gives you control over training and refuses no-referral crawlers, but it forfeits the citations and referrals that search and retrieval crawlers send back. The defensible decision is per-crawler, not all-or-nothing: block uncompensated training, price high-value retrieval with pay-per-crawl or RSL, and welcome the citation crawlers that put your content in front of users.

Which AI crawlers send traffic back?

Search-index crawlers and user-triggered fetchers: OAI-SearchBot, PerplexityBot and Claude-SearchBot cite and link sources, and ChatGPT-User or Claude-User fetch a page for a real user. Training crawlers such as GPTBot, ClaudeBot and CCBot send no referral — Cloudflare measured 1,091 OpenAI crawls and 38,065 Anthropic crawls per referral in July 2025.

Does blocking AI crawlers hurt my Google ranking?

Not if you block the right token. Google-Extended declines Gemini training and grounding and does not affect Google Search inclusion or ranking; blocking Googlebot itself removes you from Google Search and from the AI answers built on it.

Does a robots.txt block actually stop a crawler?

Only a compliant one. RFC 9309 states that its rules are not a form of access authorization, and seven of the 41 crawlers in this site's registry are documented as not honoring robots.txt; a hard block needs the CDN or a price gate such as pay-per-crawl.

What is the zero-click problem for publishers?

Users end their search on the AI answer without visiting any site: Pew Research Center measured clicks on a traditional result in 8 % of visits with an AI summary versus 15 % without (March 2025). Blocking the crawlers behind those answers removes you from the only surface many users see.

Sources

  1. Cloudflare: The crawl-to-click gap: Cloudflare data on AI bots, training, and referrals (29 August 2025), 2025. blog.cloudflare.com
  2. Cloudflare: Cloudflare just changed how AI crawlers scrape the internet-at-large (press release, 1 July 2025), 2025. cloudflare.com
  3. Pew Research Center: Google users are less likely to click on links when an AI summary appears in the results (22 July 2025), 2025. pewresearch.org
  4. OpenAI: Overview of OpenAI crawlers (GPTBot, OAI-SearchBot, ChatGPT-User), 2026. developers.openai.com
  5. Google: Google's common crawlers (Google-Extended), last updated 2026-07-14, 2026. developers.google.com
  6. IETF: RFC 9309: Robots Exclusion Protocol, 2022. rfc-editor.org
  7. AGENTS WELCOME: AI Crawler Registry — 41 crawlers and policy tokens with documented opt-out mechanism and block-vs-allow recommendation, updated 2026-07-06, 2026. agentswelcome.dev

Related: each crawler's opt-out mechanism and block-vs-allow rationale · decline training with operator opt-out tokens like Google-Extended · price access via the pay-per-crawl mechanism · license it with RSL content licensing · adoption of each standard measured over time · back to AI access economics · audit how your site treats agents.

← Access Economics · .md