# Training crawlers

> Crawlers that gather web content to train AI models.

_9 of 41 · [all The AI Crawler Registry](/crawlers)_

- [ClaudeBot](/crawlers/claudebot) — Anthropic · training. Crawls content used to train Claude. Honors robots.txt and crawl-delay.
- [GPTBot](/crawlers/gptbot) — OpenAI · training. Crawls content that may be used to train OpenAI models.
- [Meta-ExternalAgent](/crawlers/meta-externalagent) — Meta · training. Crawls content to train Meta's Llama models and AI products.
- [CCBot](/crawlers/ccbot) — Common Crawl · training. Builds the open Common Crawl corpus that many model trainers ingest downstream. Blocking CCBot blocks an upstream training-data source for the whole ecosystem.
- [Bytespider](/crawlers/bytespider) — ByteDance · training. Has a reputation for aggressive crawling and inconsistent robots.txt adherence. Rate-limit at the edge if it causes load.
- [ICC-Crawler](/crawlers/icc-crawler) — NICT (National Institute of Information and Communications Technology) · training. Crawls data to train and support AI technologies; NICT (Japan) uses the collected data for AI and may provide it to third parties, including commercial companies. Token and operator recorded in the ai.robots.txt machine-readable registry.
- [AI2Bot](/crawlers/ai2bot) — Allen Institute for AI (Ai2) · training. Collects web content to train Ai2's open language models.
- [anthropic-ai](/crawlers/anthropic-ai) — Anthropic · training. Anthropic's earlier training user-agent, widely blocked in AI robots.txt files. Anthropic's current, documented training crawler is ClaudeBot — prefer targeting ClaudeBot in new rules.
- [cohere-training-data-crawler](/crawlers/cohere-training-data-crawler) — Cohere · training. Cohere's crawler for gathering web content used to train and improve its models.
