AI crawlers

ClaudeBot: what it collects and how to control it.

Anthropic's training crawler explained: see what it collects and why blocking Anthropic's other bots may reduce your visibility in user search results.

Short answer

What is Anthropic's training crawler?

ClaudeBot is Anthropic's web crawler. According to Anthropic, it collects web content that could potentially contribute to training its generative AI models. You add its own User-agent line to robots.txt; Anthropic documents that file as the way to opt out.

In brief

Four things to know first.

  • It is for training

    Anthropic says its training crawler collects content that could potentially contribute to training its AI models.

  • Anthropic runs three bots

    Anthropic's training crawler collects web content that could contribute to model training. Claude-SearchBot improves search result quality. Claude-User visits sites when a user asks Claude a question.

  • Blocking is by robots.txt

    Anthropic says its bots honor robots.txt. Blocking IP addresses may not give a lasting opt-out.

  • Blocking Anthropic's training crawler is narrow

    It signals that your future materials should be excluded from training datasets. Search and user bots are separate.

Who runs it

What ClaudeBot does with your content

Anthropic says it uses robots to gather public web data for model development, to search the web, and to retrieve content at users' direction. Its training crawler is the one built for model development.

Anthropic says it uses separate robots so website owners get transparency and choice over each use.

  • Operator

    Anthropic, the company behind Claude.

  • Purpose

    Collecting web content that could contribute to training its generative AI models.

  • Stated aim

    Anthropic says this improves the utility and safety of its models.

  • Crawl behavior

    Anthropic says it aims to be non-intrusive and respects Crawl-delay where appropriate.

Compare

ClaudeBot vs Claude-SearchBot and Claude-User

Each Anthropic bot has its own job and its own robots.txt switch. Blocking one does not block the others.

Anthropic's training crawlerClaude-SearchBotClaude-User
Used forModel trainingImproving search resultsA user's request to Claude
What it doesCollects content that could feed trainingAnalyzes content to make search answers more relevantVisits sites when someone asks Claude a question
If you block itFuture content is excluded from training dataContent is not indexed; visibility may dropPages are not fetched; visibility may drop

Recognise it

Spot ClaudeBot in your logs

Anthropic publishes the IP addresses its crawlers use at claude.com/crawling/bots.json. It says a crawler with a source IP on that list is coming from Anthropic.

Match the Anthropic user agent name in your logs, then check the source IP against that list. A name alone proves nothing.

  1. Find the user agent

    Search your server logs for requests naming Anthropic's training crawler.
  2. Note the source IP

    Copy the IP address of each request you want to check.
  3. Check Anthropic's list

    Compare it with the IPs published at claude.com/crawling/bots.json.
  4. Report problems

    If a bot seems to malfunction, email the Anthropic crawler contact address from your site's domain.

Not sure which AI crawlers to allow?

We check which AI crawlers can reach your key pages, so blocking one bot does not cost you citations.

Control it

Block ClaudeBot in robots.txt

Anthropic says its bots respect robots.txt, so it is the right place to opt out. To block a bot from your whole site, name it and disallow everything.

Repeat the rule on every subdomain you want to opt out. Blocking IP addresses may not give a lasting opt-out, because it stops Anthropic reading your robots.txt.

Training crawler robots.txt rules, per Anthropic's documentation
Goalrobots.txt rule
Block Anthropic's training crawler site-wideUser-agent: (Anthropic's training crawler) / Disallow: /
Slow Anthropic's training crawler downName Anthropic's training crawler in User-agent, then add Crawl-delay: 1
Block search indexingUser-agent: Claude-SearchBot / Disallow: /
Block user-directed fetchesUser-agent: Claude-User / Disallow: /

The trade-off

Should you block ClaudeBot?

Blocking Anthropic's training crawler only affects training. Anthropic does not say it affects Claude's search or user-directed answers; those depend on Claude-SearchBot and Claude-User.

Anthropic warns that disabling Claude-SearchBot or Claude-User may reduce your visibility in Claude's search results. If you want citations, leave those two allowed.

BlockAllow
ClaudeBotFuture content stays out of trainingContent may contribute to training
Claude-SearchBotNot indexed; visibility may dropCan be indexed for search answers
Claude-UserPages not fetched for usersPages can be fetched on request

FAQ

ClaudeBot: common questions.

What is Anthropic's training crawler user agent?

Anthropic gives its training crawler its own user agent name. In your logs, look for requests that name that agent. Then check the source IP against the list Anthropic publishes at claude.com/crawling/bots.json before you trust it.

How do I block Anthropic's training crawler in robots.txt?

Add a User-agent line naming Anthropic's training crawler to the robots.txt file in your top-level directory. Follow it with Disallow: /. Anthropic says to repeat this on every subdomain you want to opt out. Its bots honor industry-standard robots.txt directives.

Does blocking Anthropic's training crawler remove me from Claude's answers?

Blocking Anthropic's training crawler signals that future materials should be excluded from training datasets. Anthropic describes Claude-SearchBot and Claude-User as separate bots for search and user requests. Disabling those two, not the training crawler, is what Anthropic says may reduce visibility.

What is the difference between Anthropic's training crawler and Claude-User?

Anthropic's training crawler gathers content that could contribute to model training. Claude-User visits websites when a person asks Claude a question. Anthropic says disabling Claude-User stops retrieval for user queries and may reduce your visibility for user-directed web search.

Can I slow Anthropic's training crawler down instead of blocking it?

Anthropic's training crawler can be slowed rather than blocked. [[Anthropic supports the non-standard Crawl-delay extension to robots.txt. For example, you add a User-agent line naming its training crawler, followed by Crawl-delay: 1. Anthropic says it also aims to crawl without being intrusive or disruptive.]]

Should I block Anthropic's crawler by IP address?

Blocking Anthropic's IP addresses is not recommended. Anthropic says it may not work correctly or give a lasting opt-out, because it stops its bots from reading your robots.txt file. Use robots.txt rules instead.

See which AI crawlers can reach your site.

We run your buyers' questions through ChatGPT, Claude, Perplexity and Gemini and show you the gap.