CCBot
CCBot: the Common Crawl bot, explained.
Who runs Common Crawl's crawler, what it collects, how to spot its user agent and IPs in your logs, and how to allow or block it in robots.txt.
Short answer
What is the Common Crawl bot?
CCBot is the Common Crawl bot, run by the non-profit Common Crawl. It builds an open copy of the web for researchers, companies and individuals. You identify it by the CCBot/2.0 user agent, and block it with a robots.txt Disallow.
In brief
Four things to know first.
A non-profit runs it
Common Crawl is a 501(c)(3) that provides a copy of the Internet at no cost for research and analysis.
It respects robots.txt
The crawler checks robots.txt first and fetches a page only if crawling is allowed.
You can verify it
It runs on dedicated IP ranges with reverse DNS under crawl.commoncrawl.org.
Blocking is one rule
Add a User-agent line with Common Crawl's token plus Disallow: / to robots.txt, and the crawler stops.
What it is
Who runs CCBot, and why
Common Crawl describes itself as a non-profit that provides a copy of the Internet to researchers, companies and individuals at no cost for research and analysis.
Common Crawl's crawler is built on Apache Nutch and runs on Hadoop. Common Crawl calls its dataset a sample of the web: a random subset of each site. Data sits on Amazon S3 for download.
Operator
Common Crawl, a 501(c)(3) non-profit.
Output
An open crawl dataset, stored on Amazon S3.
Coverage
A random subset of each site, not an archive of whole sites.
Your pages
It can find unlinked pages by following links from other sites.
Recognise it
CCBot user agent and IP check
Common Crawl says its crawler runs on dedicated IP ranges with reverse DNS, so you can confirm a logged request is the real one. Its ranges are also published as JSON.
| Item | Value |
|---|---|
| User agent | CCBot/2.0 (https://commoncrawl.org/faq/) |
| Reverse DNS | crawl.commoncrawl.org |
| IPv4 ranges | 3.41.188.32/29, 18.97.9.168/29, 18.97.14.80/29, 18.97.14.88/30 |
| IPv6 range | 2600:1f28:365:8000::/56 (no reverse DNS) |
| JSON list | https://index.commoncrawl.org/ccbot.json |
Behaviour
What this crawler does on your site
Common Crawl's crawler fetches pages with HTTP GET, supports HTTP/1.1 and HTTP/2, and does not execute JavaScript or use cookies. Content that needs JavaScript to appear may be missing from what it collects.
It backs off when your server returns HTTP 429 or 5xx, honors the nofollow attribute on links, and uses any Sitemap announced in robots.txt.
No JavaScript
Only the raw HTML response is read.
Adaptive back-off
Requests slow down after HTTP 429 or 5xx responses.
Redirects
Up to four redirects, or five when fetching robots.txt.
Crawl-delay
It obeys Crawl-delay in robots.txt, such as a value of 2.
Not sure which crawlers to allow?
Our free AI audit shows whether AI crawlers can reach and read the pages you want cited.
Control it
Allow or block this crawler in robots.txt
Common Crawl says it re-checks robots.txt periodically, so a change takes effect without you contacting anyone.
Open robots.txt
Edit the file at the root of your domain.Add the block rule
Add a User-agent line with Common Crawl's token, then Disallow: / on the next line.Or only slow it
Under the crawler's user-agent group, set a Crawl-delay, for example 2 for one request every 2 seconds.Optionally opt out
Common Crawl also keeps an opt-out registry you can ask to join.Check your logs
Confirm its requests stop, using the user agent and IP ranges above.
Decide
Should you block Common Crawl's bot?
Common Crawl's dataset is open, and the sources here do not say which AI models use it, so we make no claim about that. Blocking stops future collection, not data already in past crawls.
Blocking Common Crawl's bot does not touch Googlebot or the search crawlers behind AI answers. Those are separate bots with their own rules.
| Allow the Common Crawl bot | Block the Common Crawl bot | |
|---|---|---|
| Dataset | Your pages can enter open crawls | Future crawls skip your pages |
| Past data | Earlier crawls stay as they are | Earlier crawls stay as they are |
| Search crawlers | Unaffected | Unaffected |
| Server load | Tune with Crawl-delay | Requests stop |
FAQ
Common Crawl's bot: common questions.
What is the Common Crawl user agent string?
The user agent string carries version 2.0 and links to https://commoncrawl.org/faq/. Common Crawl says it replaced an older 1.0 string and the version may change, so match on the bot name, not the number.
How do I block Common Crawl's crawler?
Blocking Common Crawl's crawler takes two robots.txt lines: a User-agent line with its token, then Disallow: /. Common Crawl says it then stops crawling your website and re-checks robots.txt periodically for updates.
Does Common Crawl's crawler respect robots.txt?
Common Crawl's bot respects robots.txt. Common Crawl says it checks robots.txt first and fetches a page only if crawling is allowed. It also obeys the Crawl-delay parameter, so you can slow it down instead of blocking it.
How can I verify a request is the real Common Crawl crawler?
Verify this crawler by IP and reverse DNS. Common Crawl runs it on dedicated IP ranges with reverse DNS under crawl.commoncrawl.org, and publishes the ranges as a JSON file on index.commoncrawl.org.
Does Common Crawl's crawler execute JavaScript?
Common Crawl says its crawler does not execute JavaScript and does not use cookies. Content that only appears after scripts run may not reach its dataset, so key facts should be in the raw HTML.
Why did Common Crawl's crawler fetch pages I never linked to?
Common Crawl's crawler may reach unlinked pages by following links from other sites, Common Crawl says. It does not rely only on your own links, so a robots.txt rule is the dependable way to keep pages out.
Find out which bots can read your site.
We run your buyers' questions through ChatGPT, Claude, Perplexity and Gemini and show you the gap.