AI crawlers

Robots txt for AI crawlers: allow or block each bot.

Which user agents to name, which to allow, how to block GPTBot and other training crawlers, and what you lose in AI visibility if you block the wrong one.

Short answer

How do you set up robots.txt for AI crawlers?

Your robots txt for AI crawlers works per bot. OpenAI's settings are independent: you can disallow GPTBot to keep content out of model training and still allow OAI-SearchBot to appear in ChatGPT search. OpenAI says sites opted out of OAI-SearchBot are not shown in ChatGPT search answers.

In brief

Five things to know first.

  • Each bot has its own purpose

    OpenAI's crawlers and user agents act for its products, either automatically or when a user asks. Robots.txt can address each one separately, so you can allow one and block another.

  • Training and search are separate

    OpenAI says you can allow OAI-SearchBot to appear in search while disallowing GPTBot to opt out of training.

  • Blocking search costs visibility

    Sites opted out of OAI-SearchBot are not shown in ChatGPT search answers. Anthropic warns Claude-SearchBot blocks may reduce visibility.

  • robots.txt is a request

    Google says robots.txt cannot enforce crawler behavior. Each crawler decides whether to obey it.

  • Put a rule per bot

    Name each user agent in its own group, and repeat the rules on every subdomain you want covered.

The basics

What robots.txt does for AI bots.

Google says robots.txt tells crawlers which URLs they can access, mainly to avoid overloading your site, and cannot enforce behavior. LLMReach's free AI audit is the first step: it checks which assistants can reach you.

AI crawlers work like search crawlers: they parse HTML and follow links. Cloudflare notes they use the content as training data for machine-learning models instead of building a search index.

  • A request, not a lock

    Google says robots.txt cannot enforce crawler behavior; it is up to the crawler to obey.

  • Syntax differs by crawler

    Google warns each crawler may interpret rules differently, so use proper syntax for each bot.

  • Blocked is not hidden

    Google says a disallowed page can still be indexed if other sites link to it.

The bots

The exact user agents to name.

Use these tokens in robots.txt, taken from OpenAI's documentation. Googlebot's Image, Video and News crawlers also respond to the generic Googlebot token, so one matching token is enough for a rule to apply.

Documented AI crawler user agents and what each does, per the operator's own documentation
User agentOperatorWhat it does
GPTBotOpenAICrawls content that may train its models
OAI-SearchBotOpenAISurfaces sites in ChatGPT search
ChatGPT-UserOpenAIVisits pages when a user asks
ClaudeBotAnthropicCollects content that may train Claude
Claude-SearchBotAnthropicImproves search result quality
Claude-UserAnthropicFetches pages when a user asks

The trade-off

What blocking search crawlers costs you.

Training and search settings are independent. OpenAI says you can allow OAI-SearchBot to appear in search while disallowing GPTBot to opt out of training.

Blocking the search bots has a price. Anthropic says disabling Claude-SearchBot may reduce your visibility and accuracy in user search results.

If you blockWhat you lose
OAI-SearchBotOpenAI's search crawlerNot shown in ChatGPT search answers
Claude-SearchBotAnthropic's search crawlerMay reduce visibility in Claude search
Claude-UserAnthropic's on-request agentContent not fetched for user queries
GPTBot or ClaudeBotTraining crawlersContent excluded from training data

Not sure which bots you block?

Our free AI audit checks whether AI crawlers can reach and read your key pages, reviewed live with you.

The steps

How to block GPTBot and other crawlers.

This is how to block AI crawlers you do not want, while keeping the search bots that carry your visibility.

  1. Decide what you opt out of

    Training use and search visibility are separate choices. Decide each one before you edit anything.
  2. Block GPTBot and ClaudeBot

    Add a group for each: User-agent: GPTBot, then Disallow: /. Repeat for ClaudeBot to opt out of training.
  3. Allow the search crawlers

    Name OAI-SearchBot and GPTBot in separate groups. OAI-SearchBot surfaces sites in ChatGPT search, and OpenAI says the two settings are independent: you can allow one and disallow the other.
  4. Cover every subdomain

    Anthropic says to repeat the robots.txt block for every subdomain you want to opt out.
  5. Check and allow the IP ranges

    OpenAI recommends allowing its published IP ranges too. Allow about 24 hours for search changes to apply.

The limits

Where robots.txt falls short on its own.

Cloudflare says blocking by robots.txt depends on the operator honestly identifying itself and following RFC 9309. User agents are trivial for operators to change.

Anthropic advises against blocking its IP addresses. That stops its bots from reading your robots.txt file, so the opt-out may not hold.

  • On-request agents may ignore it

    OpenAI says robots.txt rules may not apply to ChatGPT-User, because users initiate those actions.

  • Don't block by IP alone

    Anthropic says IP blocking may not give a lasting opt-out.

  • Verify claimed bots

    Check Anthropic's published IP list at claude.com/crawling/bots.json to confirm a crawler is real.

Check yours

Check what AI crawlers can reach.

Cloudflare found only 2.98% of top sites blocked or challenged AI bots, though AI bots accessed 39% of them. Many owners have not set rules on purpose.

Technical AEO starts with access. We review which AI crawlers can read your key pages, and set robots.txt for the bots that matter in your category.

  1. Fetch your robots.txt

    Open yoursite.com/robots.txt and read each group. Note which AI user agents are named.
  2. Match rules to goals

    Confirm training bots are blocked only if you meant it, and search bots are allowed.
  3. Review your logs

    Look for the named user agents in server logs, and verify any you do not recognise.
  4. Re-check after changes

    Re-run the check after each edit. Allow time for crawlers to adjust.

FAQ

Robots.txt for AI crawlers: common questions.

How do I block GPTBot with robots txt for AI crawlers?

Blocking GPTBot takes two lines: User-agent: GPTBot and Disallow: /. OpenAI says disallowing GPTBot indicates your content should not be used in training its generative AI models. It does not affect OAI-SearchBot, which has its own rule.

Does blocking GPTBot with robots txt AI rules remove me from ChatGPT search?

Blocking GPTBot does not remove you from ChatGPT search. OpenAI says the two settings are independent: you can allow OAI-SearchBot to appear in search while disallowing GPTBot. Only opting out of OAI-SearchBot removes you from ChatGPT search answers.

What is the difference between GPTBot and OAI-SearchBot?

GPTBot crawls content that may be used to train OpenAI's models. OAI-SearchBot surfaces websites in ChatGPT's search features. They are separate user agents, so you can block one and allow the other.

Will ChatGPT-User obey my robots.txt?

ChatGPT-User may not follow robots.txt. OpenAI says it is not used for automatic crawling, and because users initiate its visits, robots.txt rules may not apply. It is also not used to decide whether content appears in Search.

How long do robots.txt changes take to apply?

For search results, OpenAI says its systems can take about 24 hours to adjust after you update robots.txt. Other operators publish no such figure, so recheck your logs after a change.

Can robots.txt stop AI scrapers completely?

Robots.txt cannot stop AI scrapers completely. Google says it cannot enforce crawler behavior, and Cloudflare notes it relies on bots honestly identifying themselves. A bot that ignores the file or spoofs its user agent needs other controls.

How do I block AI crawlers on Cloudflare?

Cloudflare offers a toggle labeled AI Scrapers and Crawlers under Security > Bots in its dashboard. Cloudflare says it blocks all AI bots and is available on the free tier.

Free AI Audit

Find out which AI crawlers can read your site.

You leave knowing where the gap is, what is causing it, and which changes would matter first, whether we work together or not.

What to expect

  1. Audit

    Before the call, we run your buyers' questions through ChatGPT, Claude, Perplexity and Gemini.

  2. Map

    On the call, we show you where competitors are cited and you are not, and what's causing it.

  3. Decide

    An honest read on whether the Citation Stack fits. If it doesn't, you'll hear it on the call.

Don't see a time that works? Email Karim directly: contact@llmreach.ai