AI crawlers
ClaudeBot: what it collects and how to control it.
Anthropic's training crawler explained: see what it collects and why blocking Anthropic's other bots may reduce your visibility in user search results.
Short answer
What is Anthropic's training crawler?
ClaudeBot is Anthropic's web crawler. According to Anthropic, it collects web content that could potentially contribute to training its generative AI models. You add its own User-agent line to robots.txt; Anthropic documents that file as the way to opt out.
In brief
Four things to know first.
It is for training
Anthropic says its training crawler collects content that could potentially contribute to training its AI models.
Anthropic runs three bots
Anthropic's training crawler collects web content that could contribute to model training. Claude-SearchBot improves search result quality. Claude-User visits sites when a user asks Claude a question.
Blocking is by robots.txt
Anthropic says its bots honor robots.txt. Blocking IP addresses may not give a lasting opt-out.
Blocking Anthropic's training crawler is narrow
It signals that your future materials should be excluded from training datasets. Search and user bots are separate.
Who runs it
What ClaudeBot does with your content
Anthropic says it uses robots to gather public web data for model development, to search the web, and to retrieve content at users' direction. Its training crawler is the one built for model development.
Anthropic says it uses separate robots so website owners get transparency and choice over each use.
Operator
Anthropic, the company behind Claude.
Purpose
Collecting web content that could contribute to training its generative AI models.
Stated aim
Anthropic says this improves the utility and safety of its models.
Crawl behavior
Anthropic says it aims to be non-intrusive and respects Crawl-delay where appropriate.
Compare
ClaudeBot vs Claude-SearchBot and Claude-User
Each Anthropic bot has its own job and its own robots.txt switch. Blocking one does not block the others.
| Anthropic's training crawler | Claude-SearchBot | Claude-User | |
|---|---|---|---|
| Used for | Model training | Improving search results | A user's request to Claude |
| What it does | Collects content that could feed training | Analyzes content to make search answers more relevant | Visits sites when someone asks Claude a question |
| If you block it | Future content is excluded from training data | Content is not indexed; visibility may drop | Pages are not fetched; visibility may drop |
Recognise it
Spot ClaudeBot in your logs
Anthropic publishes the IP addresses its crawlers use at claude.com/crawling/bots.json. It says a crawler with a source IP on that list is coming from Anthropic.
Match the Anthropic user agent name in your logs, then check the source IP against that list. A name alone proves nothing.
Find the user agent
Search your server logs for requests naming Anthropic's training crawler.Note the source IP
Copy the IP address of each request you want to check.Check Anthropic's list
Compare it with the IPs published at claude.com/crawling/bots.json.Report problems
If a bot seems to malfunction, email the Anthropic crawler contact address from your site's domain.
Not sure which AI crawlers to allow?
We check which AI crawlers can reach your key pages, so blocking one bot does not cost you citations.
Control it
Block ClaudeBot in robots.txt
Anthropic says its bots respect robots.txt, so it is the right place to opt out. To block a bot from your whole site, name it and disallow everything.
Repeat the rule on every subdomain you want to opt out. Blocking IP addresses may not give a lasting opt-out, because it stops Anthropic reading your robots.txt.
| Goal | robots.txt rule |
|---|---|
| Block Anthropic's training crawler site-wide | User-agent: (Anthropic's training crawler) / Disallow: / |
| Slow Anthropic's training crawler down | Name Anthropic's training crawler in User-agent, then add Crawl-delay: 1 |
| Block search indexing | User-agent: Claude-SearchBot / Disallow: / |
| Block user-directed fetches | User-agent: Claude-User / Disallow: / |
The trade-off
Should you block ClaudeBot?
Blocking Anthropic's training crawler only affects training. Anthropic does not say it affects Claude's search or user-directed answers; those depend on Claude-SearchBot and Claude-User.
Anthropic warns that disabling Claude-SearchBot or Claude-User may reduce your visibility in Claude's search results. If you want citations, leave those two allowed.
| Block | Allow | |
|---|---|---|
| ClaudeBot | Future content stays out of training | Content may contribute to training |
| Claude-SearchBot | Not indexed; visibility may drop | Can be indexed for search answers |
| Claude-User | Pages not fetched for users | Pages can be fetched on request |
FAQ
ClaudeBot: common questions.
What is Anthropic's training crawler user agent?
Anthropic gives its training crawler its own user agent name. In your logs, look for requests that name that agent. Then check the source IP against the list Anthropic publishes at claude.com/crawling/bots.json before you trust it.
How do I block Anthropic's training crawler in robots.txt?
Add a User-agent line naming Anthropic's training crawler to the robots.txt file in your top-level directory. Follow it with Disallow: /. Anthropic says to repeat this on every subdomain you want to opt out. Its bots honor industry-standard robots.txt directives.
Does blocking Anthropic's training crawler remove me from Claude's answers?
Blocking Anthropic's training crawler signals that future materials should be excluded from training datasets. Anthropic describes Claude-SearchBot and Claude-User as separate bots for search and user requests. Disabling those two, not the training crawler, is what Anthropic says may reduce visibility.
What is the difference between Anthropic's training crawler and Claude-User?
Anthropic's training crawler gathers content that could contribute to model training. Claude-User visits websites when a person asks Claude a question. Anthropic says disabling Claude-User stops retrieval for user queries and may reduce your visibility for user-directed web search.
Can I slow Anthropic's training crawler down instead of blocking it?
Anthropic's training crawler can be slowed rather than blocked. [[Anthropic supports the non-standard Crawl-delay extension to robots.txt. For example, you add a User-agent line naming its training crawler, followed by Crawl-delay: 1. Anthropic says it also aims to crawl without being intrusive or disruptive.]]
Should I block Anthropic's crawler by IP address?
Blocking Anthropic's IP addresses is not recommended. Anthropic says it may not work correctly or give a lasting opt-out, because it stops its bots from reading your robots.txt file. Use robots.txt rules instead.
See which AI crawlers can reach your site.
We run your buyers' questions through ChatGPT, Claude, Perplexity and Gemini and show you the gap.