
User-triggered fetchers
Agents that fetch a page in real time because a user asked — not bulk crawling.
13 of 41 The AI Crawler Registry
- Claude-User Anthropic · inference. Fetches a page in real time when a Claude user's prompt references it. User-initiated, not bulk crawling.
- ChatGPT-User OpenAI · inference. User-triggered fetch when a ChatGPT user or a GPT action requests a specific URL.
- Perplexity-User Perplexity · inference. Real-time fetch in response to a user question. Per Perplexity, user-initiated fetches are not treated as automated crawling and may ignore robots.txt — verify and rate-limit at the edge if that matters to you.
- Google-CloudVertexBot / Gemini agents Google · inference. Fetches site content on behalf of Vertex AI agents built by site owners.
- DuckAssistBot DuckDuckGo · inference. Fetches content for DuckDuckGo's AI assist answers.
- OAI-AdsBot OpenAI · ad-verification. Validates ad landing pages for OpenAI's advertising products. Listed alongside GPTBot/OAI-SearchBot/ChatGPT-User in OpenAI's bots documentation.
- Google-Agent Google · inference. User-triggered fetcher used by agents hosted on Google infrastructure to navigate the web and perform actions on a user's request (for example, Project Mariner / Gemini Agent). As a user-triggered fetcher, Google documents that it generally ignores robots.txt rules.
- MistralAI-User Mistral AI · inference. Fetches a page in real time when a Mistral (Le Chat) user's request references it. Per Mistral, the MistralAI-User token governs which sites these user-initiated requests can be made to.
- Diffbot-User Diffbot · inference. Used for requests made on behalf of human users browsing URLs through Diffbot software, as distinct from Diffbot's proactive Crawlbot. Diffbot documents both 'Diffbot' and 'Diffbot-User' as robots.txt user-agents.
- cohere-ai Cohere · inference. Retrieves data to provide responses to user-initiated prompts (Cohere products). Token and operator recorded in the ai.robots.txt machine-readable registry; the registry marks robots.txt respect as 'Unclear at this time'.
- meta-externalfetcher Meta · inference. Fetches individual links at a user's request to support Meta AI task completion. It is user-triggered, not bulk crawling.
- kagi-fetcher Kagi · inference. Fetches pages on demand for Kagi's assistant and summarizer at a user's request; not bulk crawling.
- bedrockbot Amazon · inference. Fetches web pages for Amazon Bedrock knowledge bases and web-data connectors at a customer's request; retrieval, not bulk training.