OpenAI crawler

GPTBot: what it does and how to block it.

Read OpenAI crawler requests in your logs, set robots.txt rules, and allow OAI-SearchBot so your site can appear in OpenAI's search results.

Short answer

What is OpenAI's training crawler?

GPTBot is OpenAI's web crawler. OpenAI says it crawls content that may be used to train its generative AI foundation models. You can block GPTBot in robots.txt without affecting OAI-SearchBot, the separate crawler for ChatGPT search.

In brief

Four things to know first.

  • Training, not search

    This crawler gathers content that may train OpenAI's foundation models. OAI-SearchBot handles ChatGPT search.

  • Two independent switches

    You can allow OAI-SearchBot and disallow the training crawler. OpenAI says the two settings are independent.

  • Easy to spot

    The user agent includes the crawler's name, version 1.4 and an openai.com link. OpenAI publishes its IP ranges as JSON.

  • Blocking has a clear scope

    Disallowing this crawler signals your content should not train models. It does not remove you from ChatGPT search.

What is GPTBot

What is GPTBot, and who runs it?

OpenAI runs this crawler. Its documentation says it uses crawlers and user agents to act for its products, either automatically or when a user triggers them.

This is the OpenAI crawler used for content that may be used in training OpenAI's generative AI foundation models.

OpenAI's crawlers and what each one does, per OpenAI's documentation
AgentWhat it doesrobots.txt
OpenAI's training crawlerCrawls content that may train modelsControls training use
OAI-SearchBotSurfaces sites in ChatGPT searchControls search inclusion
ChatGPT-UserVisits pages on a user's requestRules may not apply
OAI-AdsBotChecks pages submitted as adsOnly visits submitted pages

Recognise it

GPTBot user agent and IP ranges.

OpenAI gives an example user agent token for it, with version 1.4 and a link to its openai.com bot page. The version number may change.

To verify a request, compare its source IP with the JSON address list on OpenAI's bot page.

  • User agent token

    In the user-agent field of your logs, look for the crawler's name with version 1.4 and a link to OpenAI's bot page.

  • Published IP list

    OpenAI lists the crawler's addresses in a JSON file on its site. Match the source IP against it.

  • Robots.txt requests

    OpenAI may add a robots.txt marker to the user agent when fetching robots.txt, so you can tell those requests apart.

robots.txt

GPTBot robots.txt: block or allow.

OpenAI uses two robots.txt tags, one for its training crawler and one for OAI-SearchBot, so webmasters can manage how their content works with AI. Add a rule for each crawler you want to control.

  1. Open your robots.txt

    Edit the file at the root of your domain, such as yoursite.com/robots.txt.
  2. Block the training crawler

    Add a robots.txt group for GPTBot, OpenAI's training crawler, that disallows your site. OpenAI says this signals your content should not be used to train its generative AI foundation models.
  3. Keep OAI-SearchBot allowed

    Add User-agent: OAI-SearchBot and Allow: / if you want to appear in ChatGPT search results.
  4. Wait for it to apply

    For search results, OpenAI says it can take about 24 hours after a robots.txt update for its systems to adjust.
  5. Check your logs

    Confirm requests from the training crawler stop and that OAI-SearchBot still reaches your key pages.

Want AI crawlers set up right?

We review crawler access, robots.txt and rendering so the pages you want cited can be reached and read.

The trade-off

Should you block GPTBot?

Blocking OpenAI's training crawler and blocking OAI-SearchBot are different choices. Per OpenAI, sites that opt out of OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links.

OpenAI's documentation does not say blocking its training crawler affects search inclusion. Allowing both bots may let OpenAI use one crawl for both purposes.

Block the training crawler onlyBlock OAI-SearchBot too
Model trainingSignals content should not train modelsSame signal for training
ChatGPT searchPages can still be surfacedNot shown in search answers
Navigational linksCan still appearCan still appear

FAQ

GPTBot: common questions.

What is OpenAI's training crawler?

OpenAI's training crawler is its web crawler for content that may be used in training its generative AI foundation models, according to OpenAI. It is separate from OAI-SearchBot, which surfaces websites in ChatGPT search results.

How do I block OpenAI's training crawler in robots.txt?

Block this crawler in robots.txt with a User-agent line that names it, followed by Disallow: /. OpenAI says disallowing it indicates your content should not be used to train generative AI foundation models.

What is the user agent of OpenAI's training crawler?

OpenAI's docs give an example user agent token for this crawler. The token shows version 1.4 and a link to OpenAI's bot page, and the version number may change. OpenAI also publishes the crawler's IP addresses as a JSON file for verification.

Does blocking OpenAI's training crawler remove me from ChatGPT search?

Blocking OpenAI's training crawler does not control ChatGPT search inclusion; OAI-SearchBot does. OpenAI says the two settings are independent, and sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers.

Is OpenAI's training crawler the same as ChatGPT-User?

OpenAI's training crawler is not the same as ChatGPT-User. ChatGPT-User visits a page when a user asks ChatGPT or a Custom GPT a question. OpenAI says it does not crawl automatically and robots.txt rules may not apply to it.

Can I allow OAI-SearchBot and OpenAI's training crawler together?

You can allow both. OpenAI says that if both are allowed, it may use the results from just one crawl for both use cases to avoid duplicative crawling.

Check that AI crawlers can reach the pages you want cited.

We run your buyers' questions through ChatGPT, Claude, Perplexity and Gemini and show you the gap.