CCBot

CCBot: the Common Crawl bot, explained.

Who runs Common Crawl's crawler, what it collects, how to spot its user agent and IPs in your logs, and how to allow or block it in robots.txt.

Short answer

What is the Common Crawl bot?

CCBot is the Common Crawl bot, run by the non-profit Common Crawl. It builds an open copy of the web for researchers, companies and individuals. You identify it by the CCBot/2.0 user agent, and block it with a robots.txt Disallow.

In brief

Four things to know first.

  • A non-profit runs it

    Common Crawl is a 501(c)(3) that provides a copy of the Internet at no cost for research and analysis.

  • It respects robots.txt

    The crawler checks robots.txt first and fetches a page only if crawling is allowed.

  • You can verify it

    It runs on dedicated IP ranges with reverse DNS under crawl.commoncrawl.org.

  • Blocking is one rule

    Add a User-agent line with Common Crawl's token plus Disallow: / to robots.txt, and the crawler stops.

What it is

Who runs CCBot, and why

Common Crawl describes itself as a non-profit that provides a copy of the Internet to researchers, companies and individuals at no cost for research and analysis.

Common Crawl's crawler is built on Apache Nutch and runs on Hadoop. Common Crawl calls its dataset a sample of the web: a random subset of each site. Data sits on Amazon S3 for download.

  • Operator

    Common Crawl, a 501(c)(3) non-profit.

  • Output

    An open crawl dataset, stored on Amazon S3.

  • Coverage

    A random subset of each site, not an archive of whole sites.

  • Your pages

    It can find unlinked pages by following links from other sites.

Recognise it

CCBot user agent and IP check

Common Crawl says its crawler runs on dedicated IP ranges with reverse DNS, so you can confirm a logged request is the real one. Its ranges are also published as JSON.

Crawler identifiers from Common Crawl's FAQ (ranges as of 2026-08-11)
ItemValue
User agentCCBot/2.0 (https://commoncrawl.org/faq/)
Reverse DNScrawl.commoncrawl.org
IPv4 ranges3.41.188.32/29, 18.97.9.168/29, 18.97.14.80/29, 18.97.14.88/30
IPv6 range2600:1f28:365:8000::/56 (no reverse DNS)
JSON listhttps://index.commoncrawl.org/ccbot.json

Behaviour

What this crawler does on your site

Common Crawl's crawler fetches pages with HTTP GET, supports HTTP/1.1 and HTTP/2, and does not execute JavaScript or use cookies. Content that needs JavaScript to appear may be missing from what it collects.

It backs off when your server returns HTTP 429 or 5xx, honors the nofollow attribute on links, and uses any Sitemap announced in robots.txt.

  • No JavaScript

    Only the raw HTML response is read.

  • Adaptive back-off

    Requests slow down after HTTP 429 or 5xx responses.

  • Redirects

    Up to four redirects, or five when fetching robots.txt.

  • Crawl-delay

    It obeys Crawl-delay in robots.txt, such as a value of 2.

Not sure which crawlers to allow?

Our free AI audit shows whether AI crawlers can reach and read the pages you want cited.

Control it

Allow or block this crawler in robots.txt

Common Crawl says it re-checks robots.txt periodically, so a change takes effect without you contacting anyone.

  1. Open robots.txt

    Edit the file at the root of your domain.
  2. Add the block rule

    Add a User-agent line with Common Crawl's token, then Disallow: / on the next line.
  3. Or only slow it

    Under the crawler's user-agent group, set a Crawl-delay, for example 2 for one request every 2 seconds.
  4. Optionally opt out

    Common Crawl also keeps an opt-out registry you can ask to join.
  5. Check your logs

    Confirm its requests stop, using the user agent and IP ranges above.

Decide

Should you block Common Crawl's bot?

Common Crawl's dataset is open, and the sources here do not say which AI models use it, so we make no claim about that. Blocking stops future collection, not data already in past crawls.

Blocking Common Crawl's bot does not touch Googlebot or the search crawlers behind AI answers. Those are separate bots with their own rules.

Allow the Common Crawl botBlock the Common Crawl bot
DatasetYour pages can enter open crawlsFuture crawls skip your pages
Past dataEarlier crawls stay as they areEarlier crawls stay as they are
Search crawlersUnaffectedUnaffected
Server loadTune with Crawl-delayRequests stop

FAQ

Common Crawl's bot: common questions.

What is the Common Crawl user agent string?

The user agent string carries version 2.0 and links to https://commoncrawl.org/faq/. Common Crawl says it replaced an older 1.0 string and the version may change, so match on the bot name, not the number.

How do I block Common Crawl's crawler?

Blocking Common Crawl's crawler takes two robots.txt lines: a User-agent line with its token, then Disallow: /. Common Crawl says it then stops crawling your website and re-checks robots.txt periodically for updates.

Does Common Crawl's crawler respect robots.txt?

Common Crawl's bot respects robots.txt. Common Crawl says it checks robots.txt first and fetches a page only if crawling is allowed. It also obeys the Crawl-delay parameter, so you can slow it down instead of blocking it.

How can I verify a request is the real Common Crawl crawler?

Verify this crawler by IP and reverse DNS. Common Crawl runs it on dedicated IP ranges with reverse DNS under crawl.commoncrawl.org, and publishes the ranges as a JSON file on index.commoncrawl.org.

Does Common Crawl's crawler execute JavaScript?

Common Crawl says its crawler does not execute JavaScript and does not use cookies. Content that only appears after scripts run may not reach its dataset, so key facts should be in the raw HTML.

Why did Common Crawl's crawler fetch pages I never linked to?

Common Crawl's crawler may reach unlinked pages by following links from other sites, Common Crawl says. It does not rely only on your own links, so a robots.txt rule is the dependable way to keep pages out.

Find out which bots can read your site.

We run your buyers' questions through ChatGPT, Claude, Perplexity and Gemini and show you the gap.