
Training crawlers
Crawlers that gather web content to train AI models.
9 of 41 The AI Crawler Registry
- ClaudeBot Anthropic · training. Crawls content used to train Claude. Honors robots.txt and crawl-delay.
- GPTBot OpenAI · training. Crawls content that may be used to train OpenAI models.
- Meta-ExternalAgent Meta · training. Crawls content to train Meta's Llama models and AI products.
- CCBot Common Crawl · training. Builds the open Common Crawl corpus that many model trainers ingest downstream. Blocking CCBot blocks an upstream training-data source for the whole ecosystem.
- Bytespider ByteDance · training. Has a reputation for aggressive crawling and inconsistent robots.txt adherence. Rate-limit at the edge if it causes load.
- ICC-Crawler NICT (National Institute of Information and Communications Technology) · training. Crawls data to train and support AI technologies; NICT (Japan) uses the collected data for AI and may provide it to third parties, including commercial companies. Token and operator recorded in the ai.robots.txt machine-readable registry.
- AI2Bot Allen Institute for AI (Ai2) · training. Collects web content to train Ai2's open language models.
- anthropic-ai Anthropic · training. Anthropic's earlier training user-agent, widely blocked in AI robots.txt files. Anthropic's current, documented training crawler is ClaudeBot — prefer targeting ClaudeBot in new rules.
- cohere-training-data-crawler Cohere · training. Cohere's crawler for gathering web content used to train and improve its models.