Crawler guide

Googlebot: Google's Search crawler, explained.

Google Search uses two crawler types under one robots.txt token, so you cannot block mobile or desktop alone. Blocking crawling does not keep your URLs out of results.

Short answer

How do Google's AI features find your pages?

Googlebot is Google's crawler for Search. Because other crawlers often spoof its user agent, you verify it with a reverse DNS lookup or Google's published IP ranges. Blocking it in robots.txt stops crawling, but a blocked URL can still appear in search results.

In brief

Four things to know first.

  • One name, two crawlers

    Google's main crawler is the generic name for the smartphone and desktop crawlers, both run by Google Search.

  • The user agent can lie

    Other crawlers often spoof Google's crawler user agent. Check the source IP, not the header.

  • One robots.txt token

    Both crawler types obey the same token, so you cannot target the smartphone or desktop crawler alone.

  • Blocking costs you Google

    Blocking it affects Search, Discover, Images, Video and News.

Who runs it

What is Googlebot?

Google runs this crawler for Google Search. Google's documentation says it is the generic name for two crawler types: the smartphone crawler and the desktop crawler.

For most sites Google primarily indexes the mobile version. So most of its requests come from the mobile crawler, and a minority from the desktop one.

  • Purpose

    Fetches pages so Google Search can index them. It is not an AI training crawler.

  • Crawl rate

    For most sites, Google's crawler shouldn't hit you more than once every few seconds on average.

  • Size limit

    It fetches the first 2MB of a supported file type and the first 64MB of a PDF.

  • Timezone

    From US IP addresses, its timezone is Pacific Time.

Recognise it

Googlebot user agent and verification

Its user agent contains a Chrome/W.X.Y.Z string. That is a version placeholder that rises over time, so use wildcards for the version in log searches.

Google warns that other crawlers often spoof its crawler's user-agent header. Verify each request instead.

  1. Find the request

    Find the source IP of each request whose user agent says Googlebot. Google warns that other crawlers often spoof this header. Check the IP with a reverse DNS lookup or Google's published Googlebot IP ranges.
  2. Run a reverse DNS lookup

    Google says its crawler hostnames start with crawl- or geo-crawl-, followed by the IP digits, on Google's own crawler domain.
  3. Or match the IP ranges

    Compare the source IP against the IP ranges Google publishes for this crawler. Common crawlers are listed in common-crawlers.json.
  4. Block what fails

    Treat requests that fail both checks as spoofed, whatever their user agent says.

Control it

Googlebot robots.txt rules

It obeys robots.txt. Google's common crawlers always obey robots.txt rules when crawling automatically.

Google's Image, Video and News crawlers also respond to the generic token for its main crawler. You need to match only one token for a rule to apply.

Robots.txt goals for Google's main crawler and where the rule is set (source: Google Search Central)
GoalHowNote
Stop crawlingrobots.txt, token for Google's main crawlerURL can still appear in results
Stop indexingnoindex on the pagePage must stay crawlable to be read
Block all accessPassword protectionKeeps out crawlers and visitors

Not sure what your robots.txt blocks?

We check crawler access, rendering and schema so the pages you want read and cited are reachable.

The trade-off

Should you block Googlebot?

Rarely. Blocking it affects Google Search, including Discover and all Search features, plus Google Images, Video and News.

It also does not fully hide a page. Google says blocking crawling doesn't stop the URL appearing in search results. Use noindex to block indexing.

Allow Google's crawlerBlock Google's crawler
Google SearchPages can be crawled and indexedPages leave normal crawling
Other Google productsDiscover, Images, Video, News workAll are affected
Hiding a pageUse noindex or a passwordURL can still appear in results
Server loadReduce the crawl rate if neededNot the right tool for load

FAQ

Googlebot: common questions.

What is Google's main crawler?

Google's main web crawler serves Search. Google's documentation calls it the generic name for two crawler types, the smartphone crawler and the desktop crawler. It fetches pages so Google can consider them for indexing in Search.

How do I verify a request from Google's crawler?

Requests from Google's main crawler are verified by reverse DNS lookup on the source IP, or by matching the IP against Google's published crawler IP ranges. Google says the user-agent header is often spoofed by other crawlers, so the header alone proves nothing.

Can I block only the smartphone crawler?

The smartphone crawler cannot be blocked alone in robots.txt. Google says both crawler types obey the same product token, so you cannot selectively target either one using robots.txt. A rule for that token applies to both crawlers.

Does blocking Google's crawler remove a page from Google?

Blocking Google's main crawler in robots.txt does not reliably remove a page. Google says it doesn't prevent the URL from appearing in search results. Use noindex to block indexing, or password protection to block all access.

How often does Google's crawler visit my site?

For most sites, Google's crawler shouldn't access your site more than once every few seconds on average, according to Google. Short-term rates can look slightly higher. Google says you can reduce the crawl rate if your site struggles.

Does Google's crawler read my whole page?

Google's crawler reads only part of large files. For Search, it fetches the first 2MB of a supported file type and the first 64MB of a PDF. The limit applies to uncompressed data and to each referenced resource, such as CSS and JavaScript.

Make sure crawlers can reach the pages that matter.

We run your buyers' questions through ChatGPT, Claude, Perplexity and Gemini and show you the gap.